The CFO-facing infra question: a demand forecast with a fat tail, GPUs that take months to land, and three cost structures that win in different regimes. The layered commitment model that keeps you from paying for peak all year.
Plan a year of GPU fleet capacity and cost for a growing inference business. Buy, reserve, or burst?
The CFO-facing infra question: a demand forecast with a fat tail, GPUs that take months to land, and three cost structures that win in different regimes. The layered commitment model that keeps you from paying for peak all year.
Updated Aug 2026 · Grounded in real Forward Deployed Engineer interview loops and written to a senior-engineer editorial bar.
This tests whether you think like a fleet owner with a budget, not just a systems engineer. The strong answer layers committed/reserved/on-demand to match a demand curve, knows the on-demand premium is roughly 2-4x reserved, accounts for GPU lead times of months, and uses utilization (not peak) as the denominator on cost. Candidates who size to peak and rent it all on demand are off by a large multiple on cost.
No comments yet — be the first to share your approach.
