Speculative decoding
An inference technique where a small draft model proposes several tokens at once and the large model verifies them in a single forward pass, producing identical output in fewer passes.
Generation is normally one token per forward pass of the large model. Speculative decoding has a cheap draft model guess a run of tokens; the target model checks them all together and discards anything it would not have produced. The output is what the target model would have written on its own — only the number of expensive passes changes.
Where the gain comes from and where it disappears
The technique pays when verification is cheap relative to generation and the draft guesses well. Reported figures on datacentre hardware reach roughly three times throughput. On a memory-bandwidth-bound laptop running a mixture-of-experts model, the same architecture returns closer to 1.2× — neither condition holds strongly, and the speedup nearly vanishes.
Why structured output benefits most
Field names, brackets, quotes and separators are close to determined by what came before, so acceptance rates are high exactly where a user is waiting. One vendor reports 57% lower latency on function calling — a larger practical result than the throughput headline.