FDEInterviews logoFDE/Interviews

Every request you add makes the system faster and each request slower

Batching makes a model economical and makes each answer slower. Once you see that generation is memory-bandwidth bound rather than compute bound, the tuning surface becomes legible: why the first token and later tokens have different physics, and which knob to turn for which complaint.

17 MIN · PLUS

a free account unlocks the core curriculum tier · no card