FDEInterviews logo

Budget weights, resident cache and runtime memory in consistent units

Calculate a full-attention KV-cache budget with explicit units, architecture and resident-token assumptions. Compare weight precision without assuming cache precision or quality changes with it, then verify per-device allocations and runtime peaks before claiming a model fits.

30 MIN

By FDEInterviews · Updated

TL;DR: Estimate weight payload, cache for the admitted token population and runtime allocations separately, using bytes consistently. A model that loads is not yet a model that can serve the required workload; memory fit still needs per-device and peak-load verification.

Where you are. The customer has fixed hardware or a pending purchase. Before recommending a deployment, establish whether the intended architecture, context and concurrency can fit with operating margin.

Name the model and the memory being budgeted

Counting parameters is useful, but incomplete. A 70-billion-parameter payload uses 140 billion bytes at 16 bits, 70 billion at eight bits and 35 billion at four bits. Those are ideal packed weight sizes. Scales, zero points, higher-precision tensors, padding and runtime layout can change the actual allocation.

Use GB for 10⁹ bytes and GiB for 2³⁰ bytes. A 35 GB payload is about 32.60 GiB. A budget of 80 GiB contains about 85.90 GB; it is not an 80 GB budget. Record the usable memory reported for the intended environment, including what other processes reserve.

Two devices' memory cannot simply be added as if it were one addressable pool. The runtime needs a supported partitioning strategy, enough capacity on every device and a suitable interconnect. Some tensors or cache structures may be replicated. Multi-device deployment also adds communication and workspace requirements.

Derive cache size from the actual attention layout

For a standard full-attention decoder storing keys and values for every resident token, use:

KV bytes per token = 2 × layers × KV heads × head dimension × bytes per KV element
KV bytes for the batch = KV bytes per token × sum(resident tokens across sequences)

The first factor is keys plus values. Use the number of KV heads, not automatically the query-head count; grouped-query attention shares keys and values among query heads. Weight and cache precision are separate settings.

For the synthetic architecture used here, 80 layers, eight KV heads, head dimension 128 and two-byte cache elements give 327,680 bytes per token = 320 KiB. At 8,192 resident tokens, one sequence uses 2.5 GiB. The multiplication is worth doing once: 2 × 80 × 8 × 128 × 2 is 327,680 bytes per token, and 327,680 × 8,192 is 2,684,354,560 bytes, which is exactly 2.5 × 2³⁰. At 8,000 tokens it uses about 2.441 GiB instead. State the token count rather than relying on an ambiguous “8k.”

Resident tokens include the retained prompt and generated history, not only the new input. Admission planning may need to reserve space for future output or enforce a growing token budget. Multiplying every sequence by the same cap is a conservative full-reservation model when histories vary; allocating pages on demand changes occupancy but does not remove a capacity limit.

The inference chapter of How To Scale Your Model derives cache storage and its interaction with serving. Its calculations, like the worksheet below, require architectural assumptions. Sliding-window attention, latent/compressed cache designs, hybrid layers, prefix sharing and offload need different accounting. Inspect the deployed configuration rather than applying this formula to every model name.

SYNTHETIC 80 GiB BUDGET · 70B PARAMETERS · 8,192 TOKENS 8-bit weights weights 65.2 GiB 2 sequences: 5 GiB cache; about 78.6 GiB including modeled overhead. 4-bit weights weights 32.6 GiB KV cache 37.5 GiB 15 sequences fit this estimate; 16 exceed the 80 GiB budget. Same architecture and cache dtype. Weight payload changes; quality still needs evaluation. Amber: illustrative overhead = 12% of weights plus cache. Measure runtime allocations and safety reserve before making a deployment commitment.

Run the worksheet and state its approximation

This standalone Python 3 example uses 80 GiB of usable memory, 70 billion parameters and an illustrative overhead of 12% of weight-plus-cache bytes. The percentage is a sensitivity assumption, not a measured reserve or a universal runtime rule. It excludes quantization-format metadata from the weight payload.

import math

def memory_budget(params_b=70, weight_bits=4, layers=80, kv_heads=8,
                  head_dim=128, usable_gib=80, concurrency=16,
                  resident_tokens=8192, kv_bits=16, overhead=0.12):
    for name, value in {"layers": layers, "kv_heads": kv_heads,
                        "head_dim": head_dim, "resident_tokens": resident_tokens}.items():
        if type(value) is not int or value <= 0:
            raise ValueError(f"{name} must be a positive integer")
    if type(concurrency) is not int or concurrency < 0:
        raise ValueError("concurrency must be a nonnegative integer")
    for value in (params_b, usable_gib):
        if type(value) not in (int, float) or not math.isfinite(value) or value <= 0:
            raise ValueError("model size and usable memory must be positive finite numbers")
    if type(overhead) not in (int, float) or not math.isfinite(overhead) or overhead < 0:
        raise ValueError("overhead must be a nonnegative finite fraction")
    if weight_bits not in (4, 8, 16) or kv_bits not in (8, 16):
        raise ValueError("unsupported precision in this worksheet")
    gib = 2**30
    weights = params_b * 10**9 * weight_bits / 8
    kv_per_token = 2 * layers * kv_heads * head_dim * kv_bits / 8
    kv_per_sequence = kv_per_token * resident_tokens
    subtotal = weights + concurrency * kv_per_sequence
    total = subtotal * (1 + overhead)
    capacity = usable_gib * gib
    max_sequences = max(0, math.floor((capacity / (1 + overhead) - weights) / kv_per_sequence))
    return {"weights_gib": round(weights / gib, 2),
            "kv_per_sequence_gib": round(kv_per_sequence / gib, 3),
            "total_gib": round(total / gib, 2), "fits_estimate": total <= capacity,
            "max_sequences_by_memory": max_sequences}

for bits in (16, 8, 4):
    print(bits, memory_budget(weight_bits=bits))
print("four-bit, 15 sequences:", memory_budget(concurrency=15))

At the requested 16 sequences, none of the three precisions fits this 80 GiB estimate. The ideal weights occupy about 130.39, 65.19 and 32.60 GiB respectively. Cache is 40 GiB in all three cases because its precision and resident-token population did not change. The memory-only concurrency limits are zero, two and fifteen.

At four-bit weights and 15 sequences, estimated total is about 78.51 GiB. Calling that “deployable” would go beyond the worksheet: actual packed allocations, peak prefill workspace, cache bookkeeping and safety margin still need measurement. If measured usable memory is only 80 billion bytes, enter 80e9 / 2**30; the four-bit concurrency estimate falls to 13.

The interactive worksheet uses the same simplified model and explicit GiB units:

MEMORY FIT (size a deployment)
weightskv cacheoverhead118 / 80 GiB
Illustrative full-attention presets. Weights 65.19 GiB + cache 40.00 GiB, with a 12% subtotal allowance: exceeds estimate. Memory-only limit at this resident token count: 2 concurrent sequences. Measure runtime peaks and latency before setting admission limits.

Its architecture presets are illustrative parameter sets, not complete configurations for every model of that size. The slider's 8,192 tokens must include whatever retained input and output the capacity plan assumes. The memory limit is not a user-capacity or latency guarantee.

Measure overhead instead of treating a percentage as evidence

Inspect memory after loading, during the longest supported prefill (the pass that reads the whole prompt before the first output token, which is where activation memory peaks), while cache grows to the admitted limit and during the supported concurrency mix. Include runtime workspaces, graph capture, allocator reservation, fragmentation, adapters and any co-resident service. Some allocations are fixed, some scale with tokens or batch shape, and some peak during transitions.

ObservationQuestion it answersWhat to record next
Loaded model allocationActual weight/runtime footprintFormat, dtype, runtime version and device
Peak long-prompt prefillTransient activation/workspace demandPrompt length, batch and peak measurement method
Resident decode workloadCache and scheduler occupancyActive sequence lengths, cache dtype and eviction behavior
Concurrent workload transitionGrowth and admission behaviorRejections, preemptions, allocation failures and latency
Restart or reloadTemporary duplicate allocations or cold initializationPeak memory and readiness behavior

Do not subtract “allocated” from “reserved” without understanding the runtime's memory reporting. Keep units and measurement boundary in the capacity note. In a sharded deployment, repeat the accounting per device; aggregate free memory can hide an overloaded rank.

Treat precision as a tested deployment choice

Weight payload precisionIdeal 70B payloadWhat must be verified
16-bit140 GBActual memory, hardware support and required workload
8-bit70 GBQuantization format, kernels and task-quality results
4-bit35 GBMetadata/packing overhead, runtime performance and quality by segment

No precision guarantees a particular quality loss. A quantized model can match the reference on a given test and fail a different segment; the risk depends on the model, method, calibration and task. Run paired customer-relevant evaluations, including difficult and sensitive cases, before trading precision for concurrency.

QUANTIZATION (pick a precision)
65,536 levels
0.62
0.69
-0.63
-0.59
0.54
0.21
-0.65
0.13
0.89
-0.19
-0.90
0.18
0.62
-0.39
-0.34
0.75
0.31
-0.90
-0.33
0.75
0.09
-0.59
0.34
0.66
IDEAL 7B WEIGHT PAYLOAD14.0 GB
AVG ERROR0.000
This synthetic grid rounds 24 scalar values uniformly. At 16-bit grid, their mean absolute rounding error is 0.000. The ideal payload for 7 billion weights is 14.0 GB, excluding quantization metadata and runtime memory. This grid does not simulate floating-point formats or measure model quality.

The small rounding grid illustrates discretization. It does not simulate a full quantization method or measure task quality. Cache quantization is another independent choice, with its own supported formats, metadata, numerical behavior and tests. Changing weights to four bits does not imply cache changes to four bits.

Choose an overload policy from the workflow

Measure peak arrivals, admitted sequences and retained tokens rather than equating license count with concurrency. A conversation stored in a database does not necessarily occupy accelerator cache when inactive. Conversely, one long active request can consume much more cache than several short ones.

When the budget is insufficient, consider shorter retained context with evaluated retrieval, token-aware admission, bounded queues, an evaluated lower-precision or smaller model, and additional capacity. Choose among them from quality, latency, hardware and delivery constraints. Hardware can be the right early choice when a required workload cannot be reduced safely.

Queueing protects admission only when it is bounded and has a deadline. An unbounded queue does not create predictable latency. Define rejection or degradation behavior at the ceiling, including whether an output cap ends a response as incomplete. Capacity controls must not silently change the meaning of a user-facing answer.

What has to fit, in the order you count it 1 Name the usable memory GiB or GB, pick one 2 Weight payload at the chosen precision 3 Cache bytes per token from the real layout 4 Times resident tokens prompt plus history 5 Times admitted sequences not licenses sold 6 Add runtime and peaks measured, not a percentage 7 Check every device not the aggregate 8 Then say if it serves loading is not serving GB = 10^9 bytes GiB = 2^30 bytes A budget of 80 GiB holds about 85.90 GB. It is not an 80 GB budget, and the difference decides fit at the margin. 2 x layers x KV heads x head dim x bytes Use KV heads, not the query-head count: grouped-query attention shares keys and values among query heads. Weight precision and cache precision are separate settings. Inspect memory after loading, during the longest supported prefill, while the cache grows to the admitted limit, and across the supported concurrency mix. A percentage is a sensitivity assumption, not a measured reserve. Two devices' memory is not one addressable pool. Aggregate free memory can hide an overloaded rank, and some tensors or cache structures are replicated rather than split.

The spine above is the worksheet as an order of operations, with the two places the arithmetic usually goes wrong marked: the unit you chose and the head count you multiplied by.

Do this before moving on

Run the worksheet and reproduce zero/two/fifteen as the memory-only concurrency limits. Change resident tokens from 8,192 to 32,768 at four-bit weights; the per-sequence cache becomes 10 GiB and the limit falls to three. Change only KV precision to eight bits and explain why weight bytes stay fixed.

Compare 80 GiB with 80 GB and record the difference. Reject zero resident tokens, negative concurrency and non-finite overhead. Inspect the interactive controls and confirm that the four-bit 15-sequence estimate fits while 16 does not under the given assumptions.

For a real model, record its attention layout, actual weight format, cache dtype, per-device usable memory and input/output policy. Replace the percentage with measured allocations and a justified reserve. Submit a memory note that distinguishes worksheet estimates from verified serving capacity.

Go deeper

Key takeaways

  • Keep GB, GiB and exact token counts explicit throughout the estimate.
  • Cache cost follows the actual attention architecture and resident-token population.
  • Weight precision, cache precision and quality are separate decisions.
  • A memory-only fit requires runtime-peak and per-device validation before rollout.
  • Bound admission and queueing according to the workflow's latency and failure policy.

Check yourself

Answer before you look. Recalling it is what makes it stick; recognising it does not.

  1. 1In the supplied four-bit worksheet, one 8,192-token sequence uses 2.5 GiB of cache. What is the cache for 16 such sequences?

  2. 2Why does an 80 GiB budget differ from an 80 GB budget?

  3. 3The worksheet says 15 sequences fit at four-bit weights. What can you conclude?

Sign in to track which lessons you have finished.