FDEInterviews logo
MLOps & ML Engineering / 23
medium★ EssentialGoogleNVIDIAAmazon

What metrics do you autoscale inference pods on, and how do you handle cold starts?

CPU-based HPA on GPU inference never fires, the trap half of all candidates fall into within a minute. The signals that actually track load, the anatomy of a five-minute cold start, and which mitigations are worth their cost at each layer.

Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.

CPU-based HPA on GPU inference never fires, the trap half of all candidates fall into within a minute. The signals that actually track load, the anatomy of a five-minute cold start, and which mitigations are worth their cost at each layer.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 523 remaining answers · ₹2,000 / $25
UP NEXT ON YOUR JOURNEY
DISCUSSION · 0

No comments yet — be the first to share your approach.