DevOps engineer · Ahmedabad, India
I run 10+ production Kubernetes clusters on AWS, GCP, Azure, and air-gapped bare metal.
I'm the sole operator, from control-plane bootstrap through storage and the incidents. AWS Certified Solutions Architect. When something breaks, I write up what it actually turned out to be.
Recent writing
All posts-
HikariCP Pool Sizing for Postgres: Why Smaller Is Faster
How Postgres handles connections, why fewer of them run faster, and how to size HikariCP pools across many instances without exhausting max_connections.
-
etcd NOSPACE: Recovering a Kubernetes Control Plane Without kubectl
etcd hit its 2GB limit, froze all writes, and took kubectl and the API server down with it. Recovering an on-prem Kubernetes control plane using crictl over SSH.
-
When a 7-Year-Old Disk Takes Down Your Control Plane
How a 7 year old SSD caused etcd WAL fsync latency to spike above 300ms, burned the Kubernetes API error budget, and how we diagnosed it layer by layer.
Selected work
Everything else-
etcd · smartctl
Traced an error-budget-burn alert to a dying disk
KubeAPIErrorBudgetBurn on a production control plane, availability down to 90.9%. It was not load and it was not the network. smartctl found a seven-year-old SSD failing underneath etcd fsync latency. Swapped it live.
Read the write-up -
kubeadm · HAProxy · Keepalived
Bootstrapped highly available Kubernetes on bare metal
Ansible and kubeadm from scratch, with HAProxy and Keepalived holding a virtual IP in front of the API servers. Failover validated by pulling nodes on Proxmox.
-
EKS · Karpenter
Moved production from on-premise to AWS EKS
Migrated live workloads, added Karpenter for autoscaling, then spent the following weeks cutting the bill back down.
About
I like the problems where the alert is pointing at the wrong thing. A KubeAPIErrorBudgetBurn that was really a seven-year-old SSD. A control plane that was really out of etcd quota.
Most of what I know came from chasing those down and then writing them up. That is also what pulled me toward observability. In every one of those cases the fix was straightforward once the data finally said what was actually wrong, and getting to that point took far longer than it should have.
Stack
- Kubernetes
- kubeadm
- EKS
- GKE
- FluxCD
- Terraform
- Ansible
- Prometheus
- Grafana
- OpenSearch
- Fluent Bit
- Longhorn
- CrateDB
- PostgreSQL
- HAProxy
- MetalLB
- Proxmox
- Linux
- Bash
- Python
Get in touch
Open to collaborating on interesting problems. Reach me at the address below, or use the form.
kashishlakhara04@gmail.comThanks, that came through. I'll reply by email.
That didn't send. Email me directly at kashishlakhara04@gmail.com instead.