
jul 23, 2026
Your model isn't crashing, your probe is
A model server that takes five minutes to load and a liveness probe that gives it ten seconds is a crash loop waiting to happen. Probes, drains, and safe…
I do platform work at a startup. This is where I write down the things I had to figure out the hard way, so future-me at 3am doesn't have to.


SDR, a Flipper, a Pwnagotchi, a few antennas I tuned in bad lighting, and an hourglass for measuring how long a deploy actually takes.
See the lair →

jul 23, 2026
A model server that takes five minutes to load and a liveness probe that gives it ten seconds is a crash loop waiting to happen. Probes, drains, and safe…

jul 18, 2026
Same GPUs, same model, same replica count. Swap round-robin for prefix-cache-aware routing and the fleet gets 2.3x faster. The router was throwing the…

jul 16, 2026
Idle GPUs at six dollars an hour are a bonfire. Scaling to zero saves the money, but the first user back waits minutes unless you kill the cold start.

jul 11, 2026
The GPU dashboard says 92% busy and users are waiting eight seconds for the first token. Monitoring an LLM server means watching the queue, not the…
Long-form infra writing, every few weeks. No tracking pixels, no marketing sequences, no LinkedIn-isms.
Harshit Luthra. Senior SRE, infrequent essayist, occasional source of production incidents. More about me →
4M+
Daily active users
1M/min
Peak throughput
99.99%
Uptime on 95% spot
$200K
Cloud spend trimmed
What I work in
kubernetes · terraform · aws · next.js · typescript · postgres
28 posts · 11 TILs · writing here since 2018