23What metrics do you autoscale inference pods on, and how do you handle cold starts?▼medium★ EssentialGoogleNVIDIAAmazon1 replies◆ premiumCPU-based HPA on GPU inference never fires, the trap half of all candidates fall into within a minute. The signals that actually track load, the anatomy of a five-minute cold start, and which mitigations are worth their cost at each layer.Open full answer →
32Design a multi-model serving platform for LLMs with autoscaling and cold-start handling under a cost ceiling.▼hardNVIDIAGoogleAnthropic1 replies◆ premiumThe GPU-platform capstone: dozens of models, spiky traffic, a fixed monthly GPU budget, and a p99 SLO that scale-to-zero would wreck. Token-based autoscaling, the three layers of LLM cold start, and the residency tiering that makes the budget math work.Open full answer →