What is speculative decoding?
Speculative decoding pairs a large target model with a much smaller, faster draft model. The draft model guesses several tokens ahead on its own, cheaply, and the target model then checks all of those guesses in a single forward pass instead of generating them one at a time. Correct guesses are accepted for free; the first wrong one is discarded and the target generates from that point normally. The output is identical to running the target model alone; only the number of slow sequential steps changes.
Why this is worth doing
Decode is normally sequential: token N+1 depends on token N, so a model generates one token, then the next, then the next, each requiring a full pass through the network. On memory-bandwidth-bound hardware, each of those passes has to read the model's active weights from memory, and that read is what limits how fast tokens come out. Verifying several draft tokens at once still requires only one weight read, because a forward pass can score multiple candidate positions in parallel; that's the mechanism that lets speculative decoding turn several sequential steps into effectively one, when the guesses are good.
What makes a good draft model
The draft model needs to be fast enough that running it several times costs much less than one target-model step, and accurate enough that a useful fraction of its guesses get accepted. A common approach uses a smaller model from the same family (a distilled or lower-parameter version of the target) trained on similar data, since agreement between the two is more likely when they were trained similarly. Some implementations skip a separate model entirely and use the target model's own earlier layers, or an n-gram lookup table built from the current context, as a lightweight draft source instead.
Acceptance rate decides the payoff
The speedup depends entirely on how often the target model agrees with the draft. A high acceptance rate on predictable text (boilerplate code, common phrasing) can mean several tokens land per verification step, close to a multi-token speedup. A low acceptance rate, common on more creative or unpredictable text, means most draft tokens get rejected, and the wasted draft-model compute can leave speculative decoding barely faster, or in poorly tuned setups, slower, than plain decoding.
Where it fits with batching and hardware
Speculative decoding targets the same memory-bandwidth bottleneck that batching addresses, but from the opposite direction: batching amortizes one weight read across many concurrent requests' output, while speculative decoding tries to get more accepted tokens out of a single weight read for one request. That makes it most valuable for single-stream or lightly loaded serving, exactly the case the benchmarks page shows as the DGX Spark's weakest point relative to batch throughput. Running both a draft and a target model at once also means budgeting unified memory for both simultaneously, which is a real constraint worth checking against which models fit in 128 GB before assuming a pairing works.