FDEInterviews logo
MLOps & ML Engineering / 24
hardNVIDIAGoogleMicrosoft

Why does naive Kubernetes GPU scheduling strand GPUs, and how would you serve thousands of models cheaply?

The cluster shows eight free GPUs and your four-GPU pod still won't schedule, the fragmentation puzzle GPU-platform rounds open with. Bin-packing vs. spreading, MIG/MPS/time-slicing for the small-model problem, and the pod-per-model math that breaks at scale.

Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.

The cluster shows eight free GPUs and your four-GPU pod still won't schedule, the fragmentation puzzle GPU-platform rounds open with. Bin-packing vs. spreading, MIG/MPS/time-slicing for the small-model problem, and the pod-per-model math that breaks at scale.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 523 remaining answers · ₹2,000 / $25
UP NEXT ON YOUR JOURNEY
DISCUSSION · 0

No comments yet — be the first to share your approach.