Deep Dive · Systems · Two Interactive Labs

Your GPU Is Not Computing. It Is Waiting.

Generating one token from a 70B model on an H100 uses 0.34% of the chip’s arithmetic. The other 99.66% is spent waiting for memory. This is the memory wall — derived from first principles, drawn, and then taken apart optimization by optimization.

The short version — start here, the rest is the derivation

Modern accelerators can do arithmetic far faster than they can be fed. An H100 performs about 989 trillion BF16 operations per second but can only pull 3.35 terabytes per second out of its memory. Divide one by the other and you get the number that governs everything below: the chip needs ~295 arithmetic operations for every byte it reads, or the arithmetic units sit idle.

Start with a calculation anyone can check. A 70-billion-parameter model in BF16 occupies 140 GB. To generate one token, the hardware must read every one of those weights — all 140 GB — out of memory and into the arithmetic units. An H100 moves 3.35 TB/s, so that read takes 41.8 milliseconds, which caps you at 23.9 tokens per second. Not because the math is hard. Because the model had to be carried across a wire.

Now count the arithmetic. Each weight participates in one multiply and one add — two floating-point operations — and each weight cost two bytes to read. One operation per byte. The H100 is built to perform 295 operations per byte. You are asking it to do one. Everything in this post follows from that ratio, and everything people do to make inference faster is an attempt to change it.

01The model: a roof and a slope

The tool for this is the roofline model (Williams, Waterman & Patterson, 2009). It says something almost tautological and enormously clarifying. Define arithmetic intensity I as the FLOPs a kernel performs per byte it moves from memory. Then:

# attainable throughput, in FLOP/s P(I) = min( P_peak , BW × I ) # P_peak = peak arithmetic rate (FLOP/s) <- "the top" # BW = peak memory bandwidth (byte/s) <- "the slope" # I = arithmetic intensity (FLOP/byte)

Two regimes, one crossing point. When I is small you are on the slanted part: every extra byte per second of bandwidth buys you I more FLOP/s, and more arithmetic units would change nothing. When I is large you are under the flat roof: you are compute-bound and bandwidth is free. The crossing is the ridge point:

I_ridge = P_peak / BW A100 80GB : 312 TFLOP/s / 2.039 TB/s = 153 FLOP/byte H100 SXM : 989 TFLOP/s / 3.35 TB/s = 295 FLOP/byte H200 SXM : 989 TFLOP/s / 4.8 TB/s = 206 FLOP/byte

Below is that chart, live. The green line is the roof. The orange dot is your workload. The blue dots are the batch-size ladder, so you can see the whole trajectory at once.

Leave it at the defaults — 70B, BF16, batch 1 — and read the numbers. Intensity 1.0 against a ridge of 295. The dot sits on the far-left of the slope, three orders of magnitude below the roof, and the chip delivers 3.4 TFLOP/s out of 989. You bought a supercomputer and are running it as a very expensive memory controller.

A trap worth knowing about, because it doubles your answer

NVIDIA’s H100 datasheet prints BF16 tensor performance as 1,979 TFLOPS, with an asterisk: shown with sparsity. That figure requires 2:4 structured sparsity — two of every four weights exactly zero — which dense LLM GEMMs do not have. The dense number is 989. Use the headline and your ridge point comes out at 591 instead of 295, and every conclusion is off by 2×.

A neat independent check that 989 is right: FlashAttention-3 reports reaching 740 TFLOP/s, which it calls 75% utilization. 740/989 = 74.8%. Against 1,979 it would be 37%. The paper’s own percentage only makes sense against the dense figure.

02Why the wall exists: two exponentials, different exponents

This gap was not designed; it accumulated. Gholami and colleagues at Berkeley quantified it in AI and Memory Wall: over the past twenty years, peak server FLOPS scaled 3.0× every two years, while DRAM bandwidth scaled 1.6× and interconnect bandwidth 1.4×. Compounded over that window: compute up 60,000×, memory bandwidth up 100×, interconnect up 30×.

Both curves are exponential, which is why this was easy to miss, and why it is now unmissable. Their ratio is also exponential:

I_ridge(t) ∝ (3.0 / 1.6)^(t/2yr) = 1.875^(t/2yr) over 20 years: 60,000x / 100x = 600x more FLOPs demanded per byte

A machine that was balanced at a handful of operations per byte now demands three hundred. Nothing about transformer decode changed; the ground moved underneath it. And the trend is visible within a single product line — click through the three GPUs in the lab:

GPUDense BF16BandwidthRidgeBatch-1 decode uses
A100 80GB312 TFLOP/s2.04 TB/s1530.65%
H100 SXM989 TFLOP/s3.35 TB/s2950.34%
H200 SXM989 TFLOP/s4.80 TB/s2060.49%

The A100 → H100 step tripled compute and left bandwidth up only 64%, so single-stream decode efficiency got worse — 0.65% down to 0.34%. The H100 → H200 step added no compute at all and only bandwidth and capacity, and efficiency recovered to 0.49%. You are watching the wall go up and then be partially paid down, one product generation at a time.

03The derivation, properly

The one-operation-per-byte result deserves to be done carefully, because the generalization is what every optimization exploits. Take a model of N parameters at p bytes per parameter, a batch of B sequences, and k tokens produced per forward pass:

# work: each parameter does one multiply-accumulate per token FLOPs = 2 · N · B · k # traffic: the weights are read once, regardless of batch bytes = N · p # intensity I = FLOPs / bytes = 2 · B · k / p # BF16 (p = 2), one token at a time (k = 1), single stream (B = 1): I = 2 · 1 · 1 / 2 = 1 FLOP / byte

The model size N cancels entirely. A 7B model and a 405B model have the same arithmetic intensity at batch 1 — the big one is just slower in absolute terms. This is why “use a smaller model” does not fix the memory wall; it only moves you along it.

Now read that formula as a menu, because it has exactly three levers and every inference optimization pulls one of them:

LeverWhat it meansWhat pulls it
B ↑More sequences share one read of the weightsContinuous batching, higher concurrency
k ↑More tokens produced per read of the weightsSpeculative decoding, Medusa-style multi-token heads, prefill itself
p ↓Fewer bytes per weightINT8 / INT4 / FP8 weight quantization

That is the entire field in one equation. Drag the sliders in the lab above and watch: batch 4, or 4-token speculation, or int4 weights all land the dot on exactly the same spot — intensity 4 — because 2Bk/p does not care which of the three you changed. Different engineering, identical physics.

› Go deeper: prefill is the same formula with k = sequence length, which is why it feels like a different machine

Processing a prompt is not a different operation — it is the same matrix multiply with every token of the prompt available at once, so k = S, the prompt length. A 2,000-token prompt gives I = 2000, comfortably past the ridge of 295, and prefill runs compute-bound at a large fraction of peak.

So a single request crosses the roofline twice: prefill under the flat roof at high utilization, then decode collapsing onto the slope at under one percent. This is why time-to-first-token and inter-token latency behave like unrelated metrics with unrelated fixes — they are governed by opposite sides of the same chart. It is also why serving systems increasingly disaggregate the two phases onto different hardware pools: one workload wants FLOPs, the other wants bandwidth, and buying one machine for both means overpaying for whichever half you are not using.

04The second wall: what fits

Bandwidth sets your speed. Capacity sets whether you run at all, and it is the wall that bites as contexts get long. Weights are a fixed cost; the KV cache is not. Every token of every active sequence leaves behind a key and a value in every layer, held for the rest of the request:

KV_bytes = 2 · n_layers · n_kv_heads · d_head · S · B · p_kv # leading 2 = one K and one V # S = sequence length, B = concurrent sequences, p_kv = bytes/element # Llama-3-70B: 80 layers, 8 KV heads (GQA), d_head 128, BF16 per token = 2 · 80 · 8 · 128 · 2 = 327,680 bytes = 320 KiB / token

320 KiB per token sounds small. It is not, because it multiplies by both context length and user count:

At 32k context with eight concurrent users, the cache is 80 GiB — a whole H100, holding no weights. One sequence at 128k costs 40 GiB. And the control that matters most is the checkbox: turn off grouped-query attention and the same 128k sequence needs 320 GiB, more than an entire 8-GPU node after weights.

Why GQA exists, in one number

Grouped-query attention keeps all 64 query heads but shares a smaller set of key/value heads — eight, for Llama-3-70B. Queries are computed and discarded; keys and values are what you must store. So the cache shrinks by exactly the ratio n_heads / n_kv_heads = 64/8 = 8×, at a small and empirically acceptable quality cost.

Long context did not become practical because memory got bigger. It became practical because somebody noticed that the thing being cached did not need to be that wide.

05Attention has its own wall, and FlashAttention is the fix

Everything so far treated the model as weights. Attention is different: its cost is dominated not by parameters but by the S × S score matrix. Textbook attention materializes it:

S = QK⊤ / √d_head # write S² floats to HBM P = softmax(S) # read S², write S² O = PV # read S²

At 8k context that intermediate is 64 million numbers per head per layer, written to HBM and read back twice — purely to be thrown away. The arithmetic is modest; the traffic is enormous. Classic memory-bound behaviour, and it is quadratic.

FlashAttention (Dao et al., 2022) never materializes it. The matrix is tiled, each tile is loaded into on-chip SRAM, and the softmax is computed streaming using an online normalizer — a trick that predates it (Milakov & Gimelshein, 2018). The recurrence, processing key/value tile j:

m_j = max( m_{j-1} , rowmax(S_j) ) # running max α_j = exp( m_{j-1} − m_j ) # rescaling factor ℓ_j = α_j · ℓ_{j-1} + rowsum( exp(S_j − m_j) ) # running denominator O_j = α_j · O_{j-1} + exp(S_j − m_j) · V_j # running output O = O_last / ℓ_last

The factor α_j retroactively rebases everything accumulated so far from the old maximum to the new one. That is why the result is exact rather than an approximation — this is not linear attention or a sparsity heuristic, it is the same number computed with different memory traffic.

The payoff is a change in I/O complexity, which is the quantity that actually mattered all along:

standard attention : Θ( N·d + N² ) HBM accesses FlashAttention : Θ( N²·d² / M ) HBM accesses # M = SRAM size. Since d² << M, this is a large constant-factor win # on the quadratic term - the same FLOPs, far fewer bytes.

Reported results track the theory: FlashAttention-1 gave 3× on GPT-2 at 1k context and 15% end-to-end on BERT-large while reaching only 25–40% of peak; FlashAttention-2 roughly doubled that to 50–73% of theoretical max and 225 TFLOP/s (72% MFU) in end-to-end training on A100; FlashAttention-3, rebuilt around Hopper’s asynchrony with warp specialization, reaches 740 TFLOP/s in FP16 (75%) and close to 1.2 PFLOP/s in FP8.

Notice what that progression actually is. The FLOPs never changed. Three papers in a row, and every one of them is about moving fewer bytes.

06The ladder, as one motion

Now everything can be placed on one chart. Each rung below raises 2Bk/p, relieves capacity, or both — and the honest way to read any inference optimization is to ask which.

  1. Continuous batching — raises B

    The single biggest lever, because intensity is linear in it. Rather than running fixed batches, admit and retire sequences every step so the GPU never idles mid-batch. The ceiling is the ridge: past B ≈ 295 on an H100 you are compute-bound and further batching buys throughput only at the cost of latency.

  2. Weight quantization — lowers p

    In the memory-bound regime, throughput is inversely proportional to bytes per weight: BF16 → INT8 halves the traffic and doubles decode speed before any arithmetic gets faster. This is why quantization pays off far more at inference than its FLOP savings suggest. GPTQ reports ~3.25× end-to-end on A100; AWQ ~3.2×.

  3. Speculative decoding — raises k

    A small draft model proposes several tokens; the large model verifies them in one forward pass, and rejection sampling keeps the output distribution exactly that of the target model. You read the weights once and settle multiple tokens, so k > 1. Reported 2–3×. It works precisely because you were memory-bound: the verification arithmetic was free, sitting in the 99% you were not using.

  4. Grouped-query attention and MLA — relieve capacity

    GQA cuts KV by n_heads/n_kv_heads (8× for Llama-3-70B). DeepSeek’s multi-head latent attention compresses K and V into a shared low-rank latent, reported at roughly 93% reduction. Smaller cache means more concurrent sequences, which feeds back into rung 1.

  5. PagedAttention — recovers wasted capacity

    Not a smaller cache, a better-managed one. Contiguous per-request allocation fragments badly when lengths vary. vLLM manages KV in fixed blocks with an indirection table — virtual memory for attention — driving waste to near zero and reporting 2–4× throughput at equal latency. Pure systems engineering, no model change.

  6. FlashAttention — removes traffic outright

    Covered above. The quadratic intermediate never reaches HBM.

  7. Prefill/decode disaggregation — stops averaging two workloads

    One phase is compute-bound, the other bandwidth-bound. Running them on the same pool means sizing for a blend that neither wants.

07When one GPU is not enough: the third wall

Scale past a single device and a new term appears: the interconnect, which Gholami’s figures show growing slowest of all at 1.4× per two years. Each parallelism strategy trades memory for communication, and they are usually combined.

StrategyWhat it splitsWhat it costs
Tensor parallel (Megatron-LM)Individual matrices across GPUs, within a layerCollectives on the critical path — two all-reduces forward and two backward per transformer layer. Latency-bound, so it wants NVLink and stops scaling across nodes.
Pipeline parallelLayers into sequential stagesThe bubble: (P−1)/(M+P−1) for P stages and M microbatches. Cheap communication, but you need M ≫ P to keep stages busy.
Sequence / context parallel (Ring Attention)The sequence itselfKV blocks rotate around a ring, overlapped with attention compute. Enables contexts scaling with device count — the direct answer to the KV capacity wall.
ZeRO / FSDPOptimizer state, gradients, then parameters4×, 8×, then N_d× memory reduction by stage; stage 3 adds about 50% communication volume. A training-side answer.

The through-line: each one buys memory with bandwidth. Tensor parallelism gives every GPU a smaller slice of weights to read — genuinely helping decode, since the per-device read shrinks — and charges you a latency-sensitive collective on every layer. The memory wall is not removed by adding GPUs; it is converted into an interconnect wall, which is growing more slowly still.

08Recap: the whole post in six lines

If you keep nothing else:

P(I) = min( P_peak , BW × I ) # roofline: a roof and a slope I_ridge = P_peak / BW ≈ 295 on H100 # what the chip demands I = 2·B·k / p = 1 at batch 1 # what decode supplies → 1/295 = 0.34% of the chip KV = 2·L·h_kv·d·S·B·p = 320 KiB/token # the capacity wall → 80 GiB at 32k × 8 users

The first three lines say your GPU is idle almost all the time, and tell you the only three numbers that change it. The last two say that even standing still costs memory, and that the cost scales with everyone you serve. Batching, quantization, speculation, GQA, paging, FlashAttention, tensor parallelism — seven different research literatures, one shared objective: do more arithmetic per byte, or move fewer bytes.

And the reason this will not resolve itself: compute keeps compounding at 3.0× per two years and bandwidth at 1.6×. The ridge point goes up every generation. Every year you wait, the wall gets taller and the work of climbing it gets more valuable.

Takeaways

  1. The ratio that governs everything is P_peak/BW. 295 FLOPs per byte on an H100. Compute it for your hardware before optimizing anything.
  2. Batch-1 decode supplies 1 FLOP per byte. That is 0.34% of the chip, and the model size cancels out of the calculation entirely.
  3. There are exactly three levers in I = 2Bk/p: batch, tokens per pass, bytes per weight. Batch 4, 4-token speculation, and int4 weights are the same move.
  4. Capacity is a separate wall. 320 KiB per token means the KV cache outgrows the model at long context, and it scales with users.
  5. FlashAttention changed no FLOPs. Three papers, all about bytes. That is the shape of progress in this area.
  6. Multi-GPU converts the memory wall into an interconnect wall, which is growing more slowly than either compute or memory.

Sources

AI and Memory Wall (Gholami et al.) · Roofline (Williams, Waterman & Patterson, CACM 2009) · FlashAttention · FlashAttention-2 · FlashAttention-3 · Online softmax · GQA · MQA · PagedAttention / vLLM · Speculative decoding · Speculative sampling · GPTQ · AWQ · Megatron-LM · Ring Attention · ZeRO · PaLM (MFU)

On verification. Outbound fetching was blocked from the machine this was written on, so the hardware constants and reported speedups above were confirmed through search indexes rather than read off primary PDFs, and I have flagged the one place that most often goes wrong — the sparsity footnote. Every derived quantity (ridge points, intensities, KV sizes, the 600× divergence, the utilization percentages) I computed and re-checked independently, and both labs implement the same closed forms printed above. Two things I deliberately left out because I could not pin them down: the per-layer byte volume of a tensor-parallel all-reduce, and any specific decode-MFU figure.

Companion: retrieval and ANN co-design · Companion: the harness over the model

Read next: You Didn’t Pick the Best Checkpoint. · Two Towers, One Index

Cite this post
@misc{murugesan2026memorywall,
  author = {Murugesan, Sugeerth},
  title  = {Your GPU Is Not Computing. It Is Waiting.},
  year   = {2026},
  month  = {oct},
  url    = {https://sugeerth.github.io/blog/memory-wall/},
  note   = {Accessed: [date]}
}
SM
Sugeerth Murugesan Staff ML Engineer / Scientist · Intel / Intuit