Budget weights, resident cache and runtime memory in consistent units
Calculate a full-attention KV-cache budget with explicit units, architecture and resident-token assumptions. Compare weight precision without assuming cache precision or quality changes with it, then verify per-device allocations and runtime peaks before claiming a model fits.
30 MIN
By FDEInterviews · Updated
TL;DR: Estimate weight payload, cache for the admitted token population and runtime allocations separately, using bytes consistently. A model that loads is not yet a model that can serve the required workload; memory fit still needs per-device and peak-load verification.
Where you are. The customer has fixed hardware or a pending purchase. Before recommending a deployment, establish whether the intended architecture, context and concurrency can fit with operating margin.
Name the model and the memory being budgeted
Counting parameters is useful, but incomplete. A 70-billion-parameter payload uses 140 billion bytes at 16 bits, 70 billion at eight bits and 35 billion at four bits. Those are ideal packed weight sizes. Scales, zero points, higher-precision tensors, padding and runtime layout can change the actual allocation.
Use GB for 10⁹ bytes and GiB for 2³⁰ bytes. A 35 GB payload is about 32.60 GiB. A budget of 80 GiB contains about 85.90 GB; it is not an 80 GB budget. Record the usable memory reported for the intended environment, including what other processes reserve.
Two devices' memory cannot simply be added as if it were one addressable pool. The runtime needs a supported partitioning strategy, enough capacity on every device and a suitable interconnect. Some tensors or cache structures may be replicated. Multi-device deployment also adds communication and workspace requirements.
Derive cache size from the actual attention layout
For a standard full-attention decoder storing keys and values for every resident token, use:
KV bytes per token = 2 × layers × KV heads × head dimension × bytes per KV element
KV bytes for the batch = KV bytes per token × sum(resident tokens across sequences)
The first factor is keys plus values. Use the number of KV heads, not automatically the query-head count; grouped-query attention shares keys and values among query heads. Weight and cache precision are separate settings.
For the synthetic architecture used here, 80 layers, eight KV heads, head dimension 128 and two-byte cache elements give 327,680 bytes per token = 320 KiB. At 8,192 resident tokens, one sequence uses 2.5 GiB. The multiplication is worth doing once: 2 × 80 × 8 × 128 × 2 is 327,680 bytes per token, and 327,680 × 8,192 is 2,684,354,560 bytes, which is exactly 2.5 × 2³⁰. At 8,000 tokens it uses about 2.441 GiB instead. State the token count rather than relying on an ambiguous “8k.”
Resident tokens include the retained prompt and generated history, not only the new input. Admission planning may need to reserve space for future output or enforce a growing token budget. Multiplying every sequence by the same cap is a conservative full-reservation model when histories vary; allocating pages on demand changes occupancy but does not remove a capacity limit.
The inference chapter of How To Scale Your Model derives cache storage and its interaction with serving. Its calculations, like the worksheet below, require architectural assumptions. Sliding-window attention, latent/compressed cache designs, hybrid layers, prefix sharing and offload need different accounting. Inspect the deployed configuration rather than applying this formula to every model name.
Run the worksheet and state its approximation
This standalone Python 3 example uses 80 GiB of usable memory, 70 billion parameters and an illustrative overhead of 12% of weight-plus-cache bytes. The percentage is a sensitivity assumption, not a measured reserve or a universal runtime rule. It excludes quantization-format metadata from the weight payload.
import math
def memory_budget(params_b=70, weight_bits=4, layers=80, kv_heads=8,
head_dim=128, usable_gib=80, concurrency=16,
resident_tokens=8192, kv_bits=16, overhead=0.12):
for name, value in {"layers": layers, "kv_heads": kv_heads,
"head_dim": head_dim, "resident_tokens": resident_tokens}.items():
if type(value) is not int or value <= 0:
raise ValueError(f"{name} must be a positive integer")
if type(concurrency) is not int or concurrency < 0:
raise ValueError("concurrency must be a nonnegative integer")
for value in (params_b, usable_gib):
if type(value) not in (int, float) or not math.isfinite(value) or value <= 0:
raise ValueError("model size and usable memory must be positive finite numbers")
if type(overhead) not in (int, float) or not math.isfinite(overhead) or overhead < 0:
raise ValueError("overhead must be a nonnegative finite fraction")
if weight_bits not in (4, 8, 16) or kv_bits not in (8, 16):
raise ValueError("unsupported precision in this worksheet")
gib = 2**30
weights = params_b * 10**9 * weight_bits / 8
kv_per_token = 2 * layers * kv_heads * head_dim * kv_bits / 8
kv_per_sequence = kv_per_token * resident_tokens
subtotal = weights + concurrency * kv_per_sequence
total = subtotal * (1 + overhead)
capacity = usable_gib * gib
max_sequences = max(0, math.floor((capacity / (1 + overhead) - weights) / kv_per_sequence))
return {"weights_gib": round(weights / gib, 2),
"kv_per_sequence_gib": round(kv_per_sequence / gib, 3),
"total_gib": round(total / gib, 2), "fits_estimate": total <= capacity,
"max_sequences_by_memory": max_sequences}
for bits in (16, 8, 4):
print(bits, memory_budget(weight_bits=bits))
print("four-bit, 15 sequences:", memory_budget(concurrency=15))
At the requested 16 sequences, none of the three precisions fits this 80 GiB estimate. The ideal weights occupy about 130.39, 65.19 and 32.60 GiB respectively. Cache is 40 GiB in all three cases because its precision and resident-token population did not change. The memory-only concurrency limits are zero, two and fifteen.
At four-bit weights and 15 sequences, estimated total is about 78.51 GiB. Calling that “deployable” would go beyond the worksheet: actual packed allocations, peak prefill workspace, cache bookkeeping and safety margin still need measurement. If measured usable memory is only 80 billion bytes, enter 80e9 / 2**30; the four-bit concurrency estimate falls to 13.
The interactive worksheet uses the same simplified model and explicit GiB units:
Its architecture presets are illustrative parameter sets, not complete configurations for every model of that size. The slider's 8,192 tokens must include whatever retained input and output the capacity plan assumes. The memory limit is not a user-capacity or latency guarantee.
Measure overhead instead of treating a percentage as evidence
Inspect memory after loading, during the longest supported prefill (the pass that reads the whole prompt before the first output token, which is where activation memory peaks), while cache grows to the admitted limit and during the supported concurrency mix. Include runtime workspaces, graph capture, allocator reservation, fragmentation, adapters and any co-resident service. Some allocations are fixed, some scale with tokens or batch shape, and some peak during transitions.
| Observation | Question it answers | What to record next |
|---|---|---|
| Loaded model allocation | Actual weight/runtime footprint | Format, dtype, runtime version and device |
| Peak long-prompt prefill | Transient activation/workspace demand | Prompt length, batch and peak measurement method |
| Resident decode workload | Cache and scheduler occupancy | Active sequence lengths, cache dtype and eviction behavior |
| Concurrent workload transition | Growth and admission behavior | Rejections, preemptions, allocation failures and latency |
| Restart or reload | Temporary duplicate allocations or cold initialization | Peak memory and readiness behavior |
Do not subtract “allocated” from “reserved” without understanding the runtime's memory reporting. Keep units and measurement boundary in the capacity note. In a sharded deployment, repeat the accounting per device; aggregate free memory can hide an overloaded rank.
Treat precision as a tested deployment choice
| Weight payload precision | Ideal 70B payload | What must be verified |
|---|---|---|
| 16-bit | 140 GB | Actual memory, hardware support and required workload |
| 8-bit | 70 GB | Quantization format, kernels and task-quality results |
| 4-bit | 35 GB | Metadata/packing overhead, runtime performance and quality by segment |
No precision guarantees a particular quality loss. A quantized model can match the reference on a given test and fail a different segment; the risk depends on the model, method, calibration and task. Run paired customer-relevant evaluations, including difficult and sensitive cases, before trading precision for concurrency.
The small rounding grid illustrates discretization. It does not simulate a full quantization method or measure task quality. Cache quantization is another independent choice, with its own supported formats, metadata, numerical behavior and tests. Changing weights to four bits does not imply cache changes to four bits.
Choose an overload policy from the workflow
Measure peak arrivals, admitted sequences and retained tokens rather than equating license count with concurrency. A conversation stored in a database does not necessarily occupy accelerator cache when inactive. Conversely, one long active request can consume much more cache than several short ones.
When the budget is insufficient, consider shorter retained context with evaluated retrieval, token-aware admission, bounded queues, an evaluated lower-precision or smaller model, and additional capacity. Choose among them from quality, latency, hardware and delivery constraints. Hardware can be the right early choice when a required workload cannot be reduced safely.
Queueing protects admission only when it is bounded and has a deadline. An unbounded queue does not create predictable latency. Define rejection or degradation behavior at the ceiling, including whether an output cap ends a response as incomplete. Capacity controls must not silently change the meaning of a user-facing answer.
The spine above is the worksheet as an order of operations, with the two places the arithmetic usually goes wrong marked: the unit you chose and the head count you multiplied by.
Do this before moving on
Run the worksheet and reproduce zero/two/fifteen as the memory-only concurrency limits. Change resident tokens from 8,192 to 32,768 at four-bit weights; the per-sequence cache becomes 10 GiB and the limit falls to three. Change only KV precision to eight bits and explain why weight bytes stay fixed.
Compare 80 GiB with 80 GB and record the difference. Reject zero resident tokens, negative concurrency and non-finite overhead. Inspect the interactive controls and confirm that the four-bit 15-sequence estimate fits while 16 does not under the given assumptions.
For a real model, record its attention layout, actual weight format, cache dtype, per-device usable memory and input/output policy. Replace the percentage with measured allocations and a justified reserve. Submit a memory note that distinguishes worksheet estimates from verified serving capacity.
Go deeper
- GPU memory and VRAM develops the allocation and device-memory behavior behind the worksheet.
- KV cache explains retained attention state and the assumptions behind its cost.
- Quantization covers formats and evaluation beyond a scalar rounding illustration.
- Serving a 70B model on one 80GB card examines the constrained deployment decision with runtime and workload context.
- The key-value cache memory maths extends the arithmetic to mitigation and architectural choices.
- Paged attention explains allocation efficiency without treating it as unlimited cache capacity.
Key takeaways
- Keep GB, GiB and exact token counts explicit throughout the estimate.
- Cache cost follows the actual attention architecture and resident-token population.
- Weight precision, cache precision and quality are separate decisions.
- A memory-only fit requires runtime-peak and per-device validation before rollout.
- Bound admission and queueing according to the workflow's latency and failure policy.
Check yourself
Answer before you look. Recalling it is what makes it stick; recognising it does not.
1In the supplied four-bit worksheet, one 8,192-token sequence uses 2.5 GiB of cache. What is the cache for 16 such sequences?
2Why does an 80 GiB budget differ from an 80 GB budget?
3The worksheet says 15 sequences fit at four-bit weights. What can you conclude?
Sign in to track which lessons you have finished.
