Inference, KV cache and memory

On this page

What changes when the same model receives longer prompts or serves more requests? Separate prefill from decode, then account for weights and historical K/V. This article gives reproducible memory estimates while distinguishing computation, storage and measured throughput.

From a full prefix to one new position

Prefill processes a known prompt. Causal attention limits position i to its first i positions, but known input positions can be computed together in one forward pass. During decode, a sampled token becomes the next forward input: its K/V is computed and appended to history.

Future tokens do not change representations at earlier causal positions, so historical K/V can be reused. This avoids recomputing the old prefix; the new query still scores accessible keys and aggregates values. The visual counts allowed query-key pairs in one head and layer. Kernels may operate in blocks without materializing a complete score matrix. KV cache explanation

A prompt of n positions has n(n+1)/2 allowed pairs. Processing m subsequent new positions with cache adds mn+m(m+1)/2. Here sampled tokens are counted when fed into later forward passes; the first output sampled from prompt-final logits does not require another prompt pass. These are dependency counts, not FLOPs, latency or engine speed predictions.

Which positions still need computation?

Each new position makes the uncached model process the whole growing prefix; the cached model adds one query row. With n=6,m=2, initial work is 21 and new rows are 7 and 8, totaling 36. Full-prefix recomputation gives 21+28+36=85.

Preparing the visual
Which positions still need computation?

Caching avoids repeated old-query work; new queries still access historical K/V.

Space Accounting for Weights and KV Cache

Weight Storage Estimated by Total Volume

Assuming there are 35×10⁹ parameters in total, saved entirely in ideal 4-bit format, the raw weights amount to 17.5×10⁹ bytes, approximately 16.30 GiB. For FP16 raw weights, this would be 70×10⁹ bytes, approximately 65.19 GiB. Actual quantization also includes scales and grouping metadata, and some tensors retain higher precision; the bit count in a quantization format name cannot be directly taken as the final average bits per byte in the file.

If all weights reside on the same device, the device memory must be budgeted according to the total volume. Cross-device sharding, CPU offloading, or hierarchical caching are also possible, so the statement "all experts must reside on the same GPU" is not necessarily true. The cost is that when experts are selected at different positions, transmission or waiting costs may be incurred. The ability to load and the ability to serve with target latency are two separate thresholds.

KV Cache Depends on Structure and Sequence, Not Total Parameter Labels

For a simplified model that saves complete K/V for each layer, where each sequence is of equal length and uses a standard layout:

KV Bytes ≈ 2 × L × B × S × Hkv × D × b

SymbolMeaning
2Two caches for K and V
LNumber of layers using this cache structure
B, SNumber of sequences cached simultaneously, length of each sequence
Hkv, DNumber of KV heads, dimension per head
bBytes per cache element

For example, with L=32, B=1, S=8192, Hkv=8, D=128, b=2, the result is 1 GiB; after expanding the length to 65536, it becomes 8 GiB. If other parameters remain unchanged but the number of KV heads is changed to 32, the results become 4 GiB and 32 GiB respectively. This also explains why simply stating "64K context" is insufficient to estimate VRAM usage.

GQA allows multiple groups of queries to share fewer KV heads; cache volume is calculated based on the number of KV heads, not the number of query heads. GQA Paper For sliding windows, compressed caches, hybrid cyclic layers, or different layer structures, calculations should be performed layer by layer based on actual layouts; this product formula should not be mechanically applied.

from fractions import Fraction

def kv_bytes(layers, sequences, length, kv_heads, head_dim, item_bytes):
    return 2 * layers * sequences * length * kv_heads * head_dim * item_bytes

one = kv_bytes(32, 1, 8192, 8, 128, 2)
long = kv_bytes(32, 1, 65536, 8, 128, 2)
assert one == 1024**3
assert long == 8 * 1024**3
weights = Fraction(35_000_000_000 * 4, 8)
print("Ideal 4-bit Weights GiB:", round(float(weights / 1024**3), 2))
print("Example KV GiB:", one // 1024**3, long // 1024**3)
# Ideal 4-bit Weights GiB: 16.3
# Example KV GiB: 1 8

Finally, activation, operator workspace, allocator overhead, and other process usage must be added, and peak values measured under target input lengths and concurrency. The above examples do not measure a specific real-world model; specific deployment and quantization records can be found in Local LLM Deployment.

Which factor grows the KV cache?

Change one factor first, then length and concurrency together. This is a storage relationship: heads means KV heads, positions include retained history, and GiB uses 2³⁰ bytes. The model does not convert bytes into measured throughput.

Preparing the visual
Which factor grows the KV cache?

KV scales with cached positions, sequence count and KV heads; total parameter labels cannot replace those dimensions.

From Compute Volume to Actual Throughput

Activated parameters per token can help roughly estimate certain matrix computations, but they cannot be directly converted into tokens/second. Prefill for processing long prompts, step-by-step generation (decode), and high-concurrency batch processing may have different bottlenecks.

Observed PhenomenonPossible LimitationMeasurement to Compare
Long wait for the first tokenLong input processing, queuing, prefix reuse statusQueue time, prefill time, cache hit rate
Slow generation for single requestWeight or KV read, operator efficiencyLatency per step, context length, memory bandwidth
Diminishing returns when increasing batch sizeExpert hotspots, communication, workspace, or schedulingTotal throughput and tail latency per request
Can run after offloading but stuttersData movement for experts or layersInter-device transfer volume and wait times

MoE expert matrices may be small, but distribution and merging also incur overhead; multi-device expert parallelism may also require exchanging token representations. Even if two models have similar activated parameters, differences in layer count, KV structure, quantization kernels, batch size, and routing distribution can lead to significant speed differences. Therefore, the claim that "35B-A3B runs as fast as a 3B dense model" is merely a performance hypothesis awaiting verification.

Continue with:Inference-time compute, candidates and verification。