Deep Dive · Inference · Two Interactive Labs

Your Model Guesses Ahead. The Ceiling Is Exact.

Speculative decoding is the rare optimization with a closed form. A small model guesses the next few tokens, the big one checks them in a single pass, and the output distribution is unchanged. The speedup looks unbounded. It is not: it stops at 1/(1−a), and you can see the curve flatten.

Series: ML foundations (part 2 of 5)
Also in this series: The Memory Wall · Two Towers, One Index · Why Did You Rank That? · Designing a Recommender
The short version

A previous post, The Memory Wall, ended on a list of optimizations that all make the same move: settle more tokens per read of the weights. Speculative decoding is the one on that list with an exact answer, and the answer has a hard edge nobody puts on the slide.

01The trick

Generating a token requires reading every weight in the model. On an H100 running a 70B model at batch 1, that read is the whole job: the arithmetic is nearly free, and the chip spends its time waiting on memory. You are paying for a full sweep of the weights to produce one token.

So produce more than one. Have a small, cheap model — the drafter — guess the next k tokens. Then run the big model once over all of them. A single forward pass gives you the big model's opinion at every one of those positions simultaneously, because that is what a transformer forward pass does: it scores all positions at once. Walk left to right, keep the guesses the big model agrees with, stop at the first it does not, and emit the big model's own token there instead.

The part that makes this more than a heuristic is that it is exact. With the right accept/reject rule, the tokens that come out are distributed precisely as if the big model had generated them alone. You are not trading quality for speed. You are buying back the arithmetic you were already failing to use.

What you are actually paying for

Every step costs one big-model pass, plus k little-model passes, whether or not the guesses are any good. A rejected guess is not free. It is compute you already spent. That asymmetry is where the whole analysis comes from.

02The model, stated plainly

Three numbers describe the whole system:

SymbolMeaningTypical
aacceptance rate — the chance the big model agrees with any one guess0.6 – 0.9
kdraft length — how many tokens the little model guesses before being checked3 – 10
cdraft cost — one little-model pass as a fraction of one big-model pass0.02 – 0.2

A step settles a random number of tokens. If the first j guesses are accepted and the next is not, you keep j guesses and add the big model's correction: j + 1 tokens. If all k are accepted, you keep all k and the big model's next token comes free on top: k + 1. Either way, a step never settles fewer than one token — speculation cannot make you go backwards in tokens, only in time.

Take acceptance to be independent across positions with probability a. Then the number settled per step has expectation

E[tokens per step]  =  1 + a + a2 + … + ak  =  (1 − ak+1) / (1 − a)

That middle form is the one worth holding on to. The first token is certain — the big model always contributes one. The second arrives only if the first guess was accepted, with probability a. The third needs two accepts in a row, a2. Every additional token you hope for costs another factor of a, and that is the entire story.

This is a model, and here is where it bends

Independent acceptance is an idealization. In practice a depends on position (the first guess is accepted more often than the fifth), on the text (boilerplate and code accept far more readily than prose with real information in it), and on temperature. Treat a as the average over a workload, not a constant of nature. The shape of every result below survives that, because the shape comes from the geometric sum, not from the exact value.

03The ceiling

Now push k to infinity. Draft a hundred tokens ahead. A thousand. The sum 1 + a + a2 + … is geometric with ratio a < 1, so it converges, and it converges to

limk→∞ E[tokens per step]  =  1 / (1 − a)

This is the result that should change how you budget. The acceptance rate alone sets a hard upper bound on tokens per step, and no draft length gets past it:

Acceptance ak = 4k = 8k = 32Ceiling 1/(1−a)
0.51.942.002.002.00
0.72.773.203.333.33
0.83.364.335.005.00
0.94.106.139.6910.00
0.954.527.4016.3220.00

Read the a = 0.5 row. Drafting four tokens ahead already gets 1.94 of the available 2.00. Drafting eight times further buys the last 0.06 — and costs you eight times the draft compute. The row is flat because the ceiling is low, and the ceiling is low because the drafter is wrong half the time.

The ceiling binds the speedup too, not just the token count. Speedup is tokens per step divided by the cost of a step, and the cost of a step is at least one big-model pass. So

speedup  =  E[tokens] / (k·c + 1)  ≤  E[tokens]  ≤  1 / (1 − a)

Even with a free drafter — c = 0, an impossibility — an acceptance rate of 0.8 caps you at 5×. If someone promises you more, they are promising you a better drafter, not a better schedule.

04When guessing pays at all

Speculation can lose. You pay for k draft passes every step; if the guesses are usually wrong, you bought nothing with them. The condition for coming out ahead turns out to be as clean as the ceiling.

Speculation beats plain decoding if and only if a > c

The draft has to be right more often than it is expensive. Nothing else enters: not the draft length, not the model size, not the hardware.

Both directions are elementary.

If a > c, it pays. Draft a single token. The step settles 1 + a tokens on average and costs 1 + c big-model passes, so the speedup is (1 + a)/(1 + c), which exceeds 1 exactly when a > c.

If a ≤ c, nothing works. Each term of the sum past the first satisfies aj ≤ a, so the expected tokens are at most 1 + k·a. The cost is 1 + k·c. If a ≤ c then 1 + k·a ≤ 1 + k·c, so the speedup is at most 1 for every draft length. There is no clever k hiding in the tail.

I checked this numerically over 2,401 (a, c) pairs at every draft length from 1 to 400, looking for any case where the prediction and the arithmetic disagreed. There were none.

The practical reading: a drafter that costs 10% of the target model needs to be right more than 10% of the time, which is a low bar that almost any competent small model clears. The interesting question is never whether to speculate. It is how far ahead to guess — and that has an optimum, because the ceiling flattens while the cost keeps climbing.

05The first lab: watch it flatten

Two curves, two sliders. On the left, tokens settled per step against draft length, with the ceiling drawn as a dashed line. On the right, the speedup those same settings buy, with break-even marked and the optimum called out.

Start by dragging acceptance down to 0.5 and watching the left curve hit its ceiling almost immediately. Then take it to 0.95 and watch the same curve keep climbing well past k = 20 — a better drafter does not just raise the ceiling, it makes drafting further worth it. The two effects compound, which is why acceptance rate is the only number in this system worth real engineering effort.

Then drag draft cost up past the acceptance rate and watch the right-hand curve drop below break-even everywhere at once. That is the a > c boundary, and it is sharp.

At the defaults — a = 0.8, c = 0.1 — the lab reports a ceiling of 5.00 tokens per step, an optimum at k = 6, and 2.47×. Notice that the optimum settles 3.95 tokens per step, well short of the ceiling of 5.00. The economically correct draft length stops well before the ceiling, because the last tokens of headroom are the ones you are least likely to reach and you pay for the attempt every single step. Reaching 95% of the ceiling would take k = 14, more than twice as much drafting for a slower result.

Where does the optimum come from, and why is there no formula for it?

The speedup is (1 − ak+1) / ((1 − a)(k·c + 1)). Setting the derivative in k to zero gives an equation mixing a polynomial in k with ak — a transcendental equation with no closed-form root in elementary functions. In practice this does not matter at all: k is a small integer, so you evaluate 32 candidates and take the best, which is exactly what the lab does on every slider move. It is a reminder that "no closed form" and "hard" are different problems.

Sweeping acceptance at a fixed c = 0.1 gives the shape of the whole trade:

Acceptance aBest kTokens / stepSpeedupCeiling
0.521.751.46×2.00
0.632.181.67×2.50
0.742.771.98×3.33
0.863.952.47×5.00
0.9106.863.43×10.00
0.951511.204.48×20.00

The published speedups for speculative decoding cluster at 2–3×. Read backwards through this table, that range corresponds to an acceptance rate of roughly 0.70 to 0.85 — which is a reasonable description of a well-matched drafter on ordinary text, and a useful sanity check on anyone quoting a much larger number without naming their acceptance rate.

06The tail nobody budgets for

Everything above is an average, and averages are where this gets quietly misleading. Plain decoding takes exactly n passes to emit n tokens. Always. The variance is zero, which is a genuinely unusual property for a system component and one you stop noticing you depend on.

Speculation throws that away. Each step settles a random number of tokens, so finishing is a random stopping time. The mean improves. The distribution is new, and it has a right tail.

Start with the single most underappreciated fact in the whole scheme:

P(a step settles exactly one token)  =  1 − a

The very first guess is rejected with probability 1 − a, and then the step settles one token — exactly what plain decoding would have given — having paid for k draft passes on top. At a = 0.9 that is one step in ten doing no useful speculation at all. It is not a bug or a misconfiguration. It is the design operating normally.

Averaged over thousands of tokens those unlucky steps wash out. Over a short burst they do not.

07The second lab: one burst at a time

On the left, a single burst drawn step by step: green is guesses that were accepted, amber is the token the big model always contributes, faded red is guesses you paid for and threw away. On the right, the finish times of two thousand such bursts, against the dashed line where plain decoding lands every time without fail.

At the defaults — a = 0.9, k = 10, c = 0.1, a 40-token burst — the typical burst finishes in 12.0 big-model passes against plain decoding's 40. That is 3.33×, and it is a real win. But one burst in a hundred takes 20.0 or longer: 67% past the median. The distribution on the right is visibly lumpy rather than a tidy bell, because finish time is a sum of only four or five random steps and the lumps are the individual steps.

Now drag burst to the right and watch the histogram tighten:

Burst lengthp50p99p99 / p50Speedup at p50
40 tokens12.020.01.67×3.33×
100 tokens30.042.01.40×3.33×
200 tokens58.074.01.28×3.45×
400 tokens118.0138.01.17×3.39×

The speedup barely moves. The spread collapses. This is just the law of large numbers doing its job — more steps, more averaging — but it has a sharp practical edge: the benchmark and the agent are in different regimes. Speculative decoding is measured on long generations, where the tail is tame. An agent spends its life emitting short bursts: a tool name, a JSON argument, a short decision. Those bursts land in the left column of that table, where the p99 runs well past the median.

Why this bites agents specifically

A long-horizon agent is not one generation. It is hundreds of short ones chained end to end, and anything per-step compounds over a long horizon. If a step's latency has a fat right tail and you run six hundred steps, you will meet the tail many times over. The median speedup is what you report. The tail is what your timeout budget has to survive.

None of this is an argument against speculating. The median really does improve by 3× here, and that is enormous. It is an argument for measuring the thing you actually ship: if your agent's latency SLO is written at p99 and your speculation gains were measured at the mean on 1,000-token generations, those two numbers are not about the same system.

08Recap, in six lines

  1. A step settles 1 + a + a2 + … + ak tokens. Each additional token you hope for costs another factor of the acceptance rate.
  2. That sum has a ceiling of 1/(1−a), and it bounds the speedup too. At a = 0.8 you cannot exceed 5× even with a free drafter and an infinitely long draft.
  3. Guessing pays if and only if a > c. Right more often than it is expensive. Three lines to prove, no exceptions across 2,401 pairs tested.
  4. The best draft length stops well short of the ceiling, because you pay for every guess and the last tokens of headroom are the least likely to be reached.
  5. One step in 1/(1−a) settles a single token and still pays for the full draft. At a = 0.9, that is one step in ten, by design.
  6. The median gets faster; the tail gets relatively worse, and the shorter the burst the worse it gets. Agents live in the short-burst regime.

The memory wall said the chip is waiting. Speculation is the answer that says: guess while you wait, and check your guesses for free in the arithmetic you were not using. It works. It works by an amount you can write down in one line — and that line has a ceiling in it.

Sources, and what is mine

Fast Inference from Transformers via Speculative Decoding (Leviathan, Kalman & Matias, ICML 2023) · Accelerating Large Language Model Decoding with Speculative Sampling (Chen et al., DeepMind 2023) · The Memory Wall, which is why any of this is worth doing

What I could and could not verify. Direct fetches to arXiv were blocked from where this was written, so the two papers above are cited from search results, not from the PDFs. The reported speedups — 2–3× on T5-XXL with identical outputs, 2–2.5× on Chinchilla 70B — are theirs, attributed, and I have not independently reproduced them.

Everything else here is derived rather than quoted. The geometric sum, the 1/(1−a) ceiling, the a > c condition and the tail behaviour were worked out from the model stated in §02 and then checked three independent ways: against the closed form, against a direct sum over the distribution, and against a four-million-trial Monte Carlo, agreeing to within 2.2×10−3. The a > c claim was tested exhaustively over 2,401 (a, c) pairs at every draft length to 400, with no counterexample. Both labs recompute from that same closed form on every slider move, and the simulation is seeded, so the numbers in the prose are the numbers you see.

And it is a model. Constant, independent acceptance is the simplification that makes the arithmetic exact; real acceptance varies with position, content and temperature. The cost model charges k sequential draft passes plus one verify, which ignores batching effects and tree-structured drafting. The conclusions are about the shape of the trade-off, and the shape is what survives.

Cite this post
@misc{murugesan2026speculation,
  author = {Murugesan, Sugeerth},
  title  = {Your Model Guesses Ahead. The Ceiling Is Exact.},
  year   = {2026},
  url    = {https://sugeerth.github.io/blog/speculation-ceiling/}
}
SM
Sugeerth Murugesan Staff ML Engineer / Scientist · Intel / Intuit