Generating text with a large model costs one forward pass per token, and most of that pass is spent moving weights rather than doing arithmetic. Speculative decoding exploits the gap: a small draft model guesses a run of tokens, the large model checks all of them together, and whatever it would not have written is thrown away. Because the target verifies every token, greedy output is exactly what the target would have produced on its own. This is an efficiency technique, not a quality trade.
How much faster is it, really?
| Target model | Hardware | Throughput | Gain |
|---|---|---|---|
| LFM2.5-1.2B-Instruct | M4 Max MacBook | 138 to 350 tok/s | 2.54x |
| LFM2.5-2.6B | M4 Max MacBook | 61 to 139 tok/s | 2.27x |
| LFM2.5-8B-A1B | M4 Max MacBook | 90 to 106 tok/s | 1.18x |
| LFM2.5 family | H100 | reported peak | 3.18x |
| LFM2.5-VL-3B | H100, decode only | reported | up to 2.66x |
| LFM2.5-VL-3B | M5 Max, decode only | reported | 2.30x to 3.13x |
The row worth staring at is the 8B mixture-of-experts model on the MacBook: 1.18x. The same drafter design on an H100 returns 2.54x on a comparable target. Nothing about the implementation changed between those two numbers.
Why does the same technique give 1.18x and 3.18x?
Two conditions have to hold together. Verification must be cheap relative to generation, which is a property of the hardware: it holds where the machine has compute to spare and is waiting on memory. And the draft model must guess well, which is a property of the text being produced. On a memory-bandwidth-bound device running a model with around one billion active parameters, neither holds strongly, and the technique nearly disappears.
This is why a published speedup is close to useless as a purchasing input. The number belongs to a target model, a drafter, a piece of hardware and a traffic mix. Anyone already serving the model can measure it against their own traffic in an afternoon, and that is the only measurement worth acting on.
Why is the end-to-end gain smaller than the decode gain?
Because only decode accelerates. Prefill — processing the prompt — runs exactly as before, and in a vision-language model the image passes through an encoder first and arrives as hundreds of visual tokens that also have to be prefilled. The vision drafter we covered reports decode 2.30x to 3.13x faster on an M5 Max while end-to-end latency improves 1.56x to 2.62x; on an M3 Ultra under llama.cpp the same split is 1.57x to 2.14x against 1.30x to 1.77x.
The authors name the reason plainly, and it is the oldest one in performance engineering: the overall speedup is capped by the part of the workload that was not accelerated. On edge hardware prefill takes a larger share of wall time than in a datacentre, so the same decode gain shows up smaller at the user.
Which workloads gain most?
Structured output, by a wide margin. Field names, brackets, quotes and separators are close to determined by what came before, so the draft model's proposals are accepted far more often. Liquid AI measured 57% lower latency on function calling with a 2.6B target — on exactly the path where a user is sitting still waiting for a tool call to return.
- Function calling and JSON output: high acceptance, and the latency is directly visible to a user.
- Code generation: repetitive syntax, long runs the drafter gets right.
- Open prose: lower acceptance, and the gain shrinks accordingly.
What does it cost?
Memory, and not much of it. The drafters in these releases are small and deliberately plain: around 295.7M parameters for a 1.2B target and 327.7M for the 2.6B and 8B models, attention-only with five decoder layers. The vision drafter is about 280M parameters on a 3B target, which is 8.9% added to what is deployed, built as four layers with a block size of nine after ablations across three, four and five.
There is a second cost that does not appear in any benchmark table: a second model to version, evaluate and ship. A drafter trained against one target is not valid against another, and nothing warns you when the pairing drifts.
Where it has already changed a comparison
In hardware evaluation the technique has quietly become part of the baseline. SemiAnalysis measured OpenAI's Jalapeno inference chip ahead of every other part on throughput per megawatt without using speculative decoding, while the chips it was compared against were using it. A benchmark that does not say whether speculation was on is not comparing the same thing twice.
What to establish before adopting it
- Measure decode and end-to-end separately. If prefill dominates your traffic, the decode figure will flatter the result.
- Measure on your own traffic mix, not the published one: acceptance rate is a property of what your users ask for.
- Check whether your serving stack supports it for your target — these releases shipped with specific builds of llama.cpp, MLX-VLM and SGLang rather than with the mainline.
- Confirm the output guarantee holds in your configuration: it is exact under greedy decoding because the target verifies every token.
- Decide who owns the drafter when the target model is upgraded.
Questions
- Does speculative decoding change the model's answers?
- No. The target model verifies every proposed token and discards anything it would not have produced, so greedy output is identical to running the target alone. Only the number of expensive forward passes changes.
- Why did speculative decoding barely help on my laptop?
- Most likely because the machine is memory-bandwidth-bound and the model has few active parameters, so verification is not cheap relative to generation. We have covered a case where an 8B mixture-of-experts target returned 1.18x on an M4 Max MacBook while comparable targets reached 2.54x on datacentre hardware.
- How large does the draft model need to be?
- Small. In the releases we have covered the drafters are around 280M to 330M parameters against targets of 1.2B to 8B, attention-only with four or five decoder layers. One of them added 8.9% to the deployed parameter count.
- Which workloads should be tested first?
- Anything structured. Field names, brackets and separators are close to determined by what precedes them, so acceptance is high; one vendor measured 57% lower latency on function calling, which is where a user is actually waiting.