The Attention Gap
Two numbers that look like one — and the 8× they hide in every KV cache estimate
Published: 2026-10-04 | jslet Research | 13 min read
📑 In This Briefing
Ask a well-built KV cache calculator what a 70B model needs at 8K context and it will tell you about 21 GB per sequence. For almost every 70B model anyone has deployed since 2023, the true figure is 2.62 GB. The gap is a factor of eight, and it is not overhead, fragmentation, or a disagreement about batch sizes. It is one substitution: a head dimension computed against KV heads where it should be computed against attention heads.
The substitution is easy to make because the algebra appears to close. Write the per-token cache as 2 × layers × kv_heads × (d_model / kv_heads) × bytes and the kv_heads terms divide out, leaving 2 × layers × d_model × bytes — a formula that is correct, tidy, and describes a model with no grouped-query attention at all. Every model released in the last three years uses grouped-query attention. The distance between those two facts is what this briefing is about.
It is the written companion to the KV Cache Context Window calculator, which now computes both forms side by side so the difference is something you can see rather than something you have to trust.
The Formula That Cancels Itself
The cache footprint per token, in bytes, is:
Six terms, and every one of them is load-bearing. The 2 covers keys and values, because both are stored. layers multiplies because each transformer block keeps its own cache — a 126-layer model pays the per-token cost 126 times. kv_heads × head_dim is the width of the key and value projection. bytes_per_element is the storage precision, 2 for FP16. And context × batch is what makes the cache a variable cost rather than a fixed one: it grows with every token you hold and every sequence you serve at once.
The error enters at head_dim. Head dimension is d_model divided by the number of attention heads. Write it as d_model / kv_heads instead and the expression becomes 2 × layers × kv_heads × (d_model / kv_heads) × bytes…, where kv_heads sits in both numerator and denominator and cancels cleanly. What survives is 2 × layers × d_model × bytes — the textbook multi-head attention figure, in which the KV head count has no effect whatsoever.
This is the part worth pausing on. The error does not underestimate the grouped-query saving; it deletes the term entirely. Any calculator or spreadsheet that accepts the cancelling form will return the same number whether you tell it the model has 8 KV heads or 64, because after the cancellation it never asks.
| Expression | What it actually computes | 70B, 8K, batch 1, FP16 |
|---|---|---|
2 × L × kv_heads × (d_model / kv_heads) × 2 × ctx | Multi-head attention — the KV head count cancels | 20.97 GB |
2 × L × kv_heads × (d_model / n_heads) × 2 × ctx | Grouped-query attention — 8 KV heads of width 128 | 2.62 GB |
The two rows differ by the query-to-KV head ratio. On the 70B that ratio is 8. On a model that has moved all the way to multi-query attention it is 64. And on a model with a wide head dimension and a large KV head count it can be 1 — the one case where the cancelling form happens to be right.
Two Numbers That Look Like One
Attention geometry is a spectrum, not a binary, and that is precisely why the confusion survives review. Three layouts matter:
| Layout | Query heads | KV heads | Cache vs MHA | 70B, 8K, FP16 |
|---|---|---|---|---|
| Multi-head (MHA) | 64 | 64 | 1.00× | 20.97 GB |
| Grouped-query (GQA-8) | 64 | 8 | 0.125× | 2.62 GB |
| Multi-query (MQA) | 64 | 1 | 0.016× | 0.33 GB |
In multi-head attention every query head owns a key head and a value head, so the cache scales with the head count. In grouped-query attention, groups of query heads share a single key/value head — this is what Llama-2-70B introduced with its 64 query heads and 8 KV heads, and it became the default for essentially everything that followed. In multi-query attention all query heads share one key/value head, the maximum compression available before you change the attention mechanism itself.
Both GQA and MHA feed kv_heads into the formula, which is what makes them look interchangeable. An estimator that is never told the query-head count cannot tell them apart, and will silently answer for the wrong one. The failure is worst exactly where it matters most: the wider the query-to-KV ratio, the larger the error, so a model that has invested the most engineering in shrinking its cache is the one whose cache gets overstated the most.
The information needed to get this right is not secret. Model config files report both fields — num_attention_heads and num_key_value_heads — and they are equal only for older MHA checkpoints. The gap between them is the grouped-query factor, and it is the first thing a sizing calculation should read.
What the 8× Is Worth
The 70B class is the useful reference point because it is the largest model that still fits on a single 80 GB card at reduced precision, which makes the cache the term that decides the deployment rather than the weights. One sequence, FP16 cache, 80 layers, 8,192 hidden width, 64 attention heads and 8 KV heads:
| Context length | KV cache, batch 1 | Weights (FP16) | Cache as % of weights |
|---|---|---|---|
| 4,000 tokens | 1.31 GB | 140 GB | 0.9% |
| 8,000 tokens | 2.62 GB | 140 GB | 1.9% |
| 32,000 tokens | 10.49 GB | 140 GB | 7.5% |
| 128,000 tokens | 41.94 GB | 140 GB | 30.0% |
Two things follow. The first is that the cache is linear in context length, not exponential — the 16× context increase from 8K to 128K produces a 16× cache increase. It only looks explosive because it is multiplied by concurrency, which is the subject of the next-but-one section. The second is that on a single sequence, even at 128K, the cache is still under a third of the weights. Anyone sizing from the cancelling formula would conclude the opposite: at 21 GB per 8K sequence they would see the cache overtake the weights somewhere around 53K context, and plan a multi-GPU tensor-parallel layout that the real workload never requires.
The 8× is also not a constant. It is the query-to-KV ratio, and it varies by model — as does the per-token cache itself, which depends on layer count and KV head count rather than on parameter count:
| Model | Layers | KV heads | Per token, FP16 | 8K, batch 1 | 128K, batch 1 |
|---|---|---|---|---|---|
| Llama 4 Scout 8B | 32 | 8 | 128 KiB | 1.05 GB | 16.78 GB |
| Llama 4 Maverick 70B | 80 | 8 | 320 KiB | 2.62 GB | 41.94 GB |
| DeepSeek-V3 671B (MoE) | 61 | 8 | 244 KiB | 2.00 GB | 31.98 GB |
| Mistral Large 2 123B | 88 | 8 | 352 KiB | 2.88 GB | 46.14 GB |
| Gemma 3 27B | 62 | 16 | 496 KiB | 4.06 GB | 65.01 GB |
| Llama 4 Behemoth 405B | 126 | 8 | 504 KiB | 4.13 GB | 66.06 GB |
Read the highlighted comparison against the row below it. A 27B model with 16 KV heads carries a per-token cache a little over 1.5× larger than a 70B model with 8 — 496 KiB against 320 KiB. Parameter count is not a proxy for cache footprint, and neither is weight size. A 671B mixture-of-experts model, the largest thing in this table by parameter count, has a smaller per-token cache than a 123B dense model, because its 61 layers put it below Mistral Large 2's 88.
Precision Is a Separate Dial
There is a second error in this area, and this one runs the other way. Weight quantization and cache quantization are different settings, and treating them as one understates the cache by exactly the same class of factor it is commonly overstated by.
| Component | FP16 weights | INT4 weights |
|---|---|---|
| 70B weights | 140 GB | 35 GB |
| KV cache, 8K, batch 1 | 2.62 GB | 2.62 GB |
Quantizing weights to INT4 divides the weight footprint by four and leaves the cache untouched, because in essentially every serving stack the cache stays at FP16 or BF16 unless you configure it otherwise. vLLM, TensorRT-LLM and SGLang all ship FP16 KV by default, and the flag that changes it is a KV cache flag, not a weight flag. Sizing a deployment as though 4-bit weights implied a 4-bit cache understates the cache by 4× — the mirror image of the head-dimension error, and just as capable of producing an out-of-memory surprise under load.
Cache precision is a real lever, and a powerful one: INT8 or FP8 keys and values halve the cache, which on the 70B is 21 GB back at 128K context. But it is a separate decision with a separate quality trade, and it should be made deliberately rather than assumed as a side effect of the weight format. The rule is short: read the cache precision from the serving configuration, not from the model name.
Where the Cache Actually Bites
The single-sequence view is the one that makes people think long context is not a memory problem. The concurrent view is the one that actually decides hardware. Same 70B, same FP16 cache, same 128K context — only the batch size changes:
| Concurrent sequences at 128K | KV cache | vs FP16 weights (140 GB) |
|---|---|---|
| 1 | 41.94 GB | 0.3× |
| 8 | 335.5 GB | 2.4× |
| 32 | 1,342 GB | 9.6× |
At one sequence the cache hides under the weights. At eight it has overtaken them. At thirty-two it is nearly ten times the model, and the number that decides your GPU count is no longer the parameter size at all. This is the honest version of the claim that long context is expensive: not that a single 128K request is a memory disaster — it is 42 GB, which is containable — but that the thing you are actually selling is dozens of concurrent long conversations, and each one carries its own 42 GB.
The consequence for layout is direct. When the cache term dominates, tensor parallelism has to cover the cache, not just the weights, and the group size is set by how many concurrent sequences of the target context length you need to hold. That is a different calculation from the one that sizes the weights, and it produces a larger fleet — which is why a deployment planned from parameter count arrives at a GPU count that is too low, while one planned from the cancelling cache formula arrives at a count that is too high for the wrong reason. The fix in both directions is the same: compute the cache with the right head dimension, then multiply by the concurrency you intend to serve.
A Four-Number Sizing Rule
The whole calculation reduces to four quantities. Get them right and the rest is multiplication.
- Head dimension = d_model ÷ attention heads. Not KV heads. On a 70B that is 8,192 ÷ 64 = 128. If you only have the KV head count, you cannot compute the cache — go and find the query head count, because the two differ on every modern model.
- Per-token cache = 2 × layers × kv_heads × head_dim × bytes. With FP16 cache that is 2 bytes per element. For the 70B: 2 × 80 × 8 × 128 × 2 = 320 KiB per token, or 0.328 GB per 1,000 tokens.
- Multiply by the context you will actually hold, then by concurrency. The product is the cache, and it is the term that moves fastest because both factors are under your product's control rather than yours. A cache sized for 8K single-sequence traffic is 16× too small for 128K at eight concurrent — a 128× miss from two decisions that each look modest.
- Add the weights and check the sum against the card — then let the larger term set the layout. Weights are set by parameters and precision; cache is set by attention geometry, context and concurrency. When the cache wins, the tensor-parallel group is sized by the cache.
And one judgement call that sits outside the four: state the cache precision explicitly in the sizing record, because it is a 4× lever that defaults silently. FP16 is a defensible default; it just should not be an unexamined one.
Steps one and two are what the KV cache calculator computes directly, with separate inputs for attention heads and KV heads. Step four's weight side is the FP16 VRAM estimator or the INT4 equivalent, and the question of whether the resulting fleet fits at all is the fit matrix.
Frequently Asked Questions
How much VRAM does a 70B model's KV cache need?
For a 70B with 80 layers, 8,192 hidden width, 64 attention heads and 8 KV heads, at FP16 cache and one sequence: about 2.62 GB at 8K context, 10.49 GB at 32K and 41.94 GB at 128K. Per token it is 320 KiB. Multiplying by concurrency is where the numbers become large — 32 concurrent sequences at 128K is 1,342 GB of cache against 140 GB of weights.
Why is my KV cache estimate eight times too big?
Almost certainly because the head dimension was computed as d_model divided by the KV head count instead of the attention head count. Written that way, the KV head count appears in both numerator and denominator and cancels, leaving the full multi-head attention formula — so the grouped-query saving disappears entirely. The ratio between the two results equals query heads divided by KV heads, which is 8 on a typical 70B.
Does quantizing the weights to INT4 shrink the KV cache?
No. Weight precision and cache precision are independent settings. INT4 weights divide the 70B weight footprint from 140 GB to 35 GB and leave the FP16 cache exactly where it was, because serving stacks keep the cache at FP16 or BF16 unless a separate KV cache flag changes it. Assuming a 4-bit cache because the weights are 4-bit understates it by 4×.
At what context length does the KV cache exceed the model weights?
For one sequence, it does not — at 128K the 70B's cache is 41.94 GB against 140 GB of FP16 weights, about 30%. The comparison flips with concurrency: eight concurrent 128K sequences carry 335.5 GB, or 2.4× the weights, and thirty-two carry 1,342 GB, or 9.6×. The question is therefore about batch size as much as about context length.
Methodology & Disclosure
Cache sizes are computed as 2 × layers × kv_heads × head_dim × bytes_per_element × context × batch, where head_dim = d_model / attention_heads. Model geometry — layer counts, KV head counts and head dimensions — is taken from each model's published architecture configuration, not estimated from parameter counts. Context lengths use decimal K (8K = 8,000 tokens, 128K = 128,000) and sizes are reported in decimal GB, consistent with the site's machine-readable numbers file. Framework overhead, allocator fragmentation, PagedAttention block padding and attention kernel workspace are excluded, so these figures are lower bounds on real allocation. The 70B reference shape is 80 layers, 8,192 hidden width, 64 attention heads, 8 KV heads, FP16 cache. Cache precision is assumed FP16 throughout unless a section says otherwise; INT8 and FP8 caches halve the figures. This site takes no sponsorship, no affiliate commission and no vendor payment of any kind; the vendors named here have no relationship with jslet and did not review this briefing.
References & Further Reading
- Ainslie et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — the attention layout and the query-to-KV head ratio used throughout. arXiv:2305.13245
- Shazeer. Fast Transformer Decoding: One Write-Head is All You Need — multi-query attention, the limiting case of the same geometry. arXiv:1911.02150
- Meta AI. Llama 2, Llama 3 and Llama 4 model cards and configuration files — layer counts, attention head counts and KV head counts per size class. ai.meta.com
- vLLM documentation. PagedAttention, KV cache memory management and the
kv_cache_dtypesetting that decouples cache precision from weight precision. docs.vllm.ai - NVIDIA TensorRT-LLM documentation. KV cache quantization, paged KV memory and multi-GPU deployment guidance. nvidia.github.io/TensorRT-LLM
- Related: KV Cache Context Window Calculator · The VRAM Ceiling: Why 70B Is a Topology Choice · GPU Model Fit Matrix · Parameters to VRAM (FP16) · The Reverse Sizing Problem: From QPS to Silicon
📜 Copyright & Attribution
© 2026 jslet Research. This article is an original work published on jslet (jslet.com). All rights reserved.
Sharing & Reprinting: You may share excerpts (up to 200 words) with a mandatory, do-follow link back to the original article URL. Full reproduction, translation, or adaptation requires prior written permission. Contact: support@jslet.com.
Estimation Disclaimer: Cache figures are model-based estimates from published architecture parameters, not measurements. Real allocation varies with framework, allocator behaviour, block padding and kernel selection, and is typically higher than the figures here. Validate against measured GPU memory on your target stack before committing to hardware.