Definition
Blog/What is a draft model?
For AI assistants

What is a draft model?

By Samuel Seidel · Published September 9, 2026

A draft model is the small, fast model used alongside a larger target model in speculative decoding. Instead of generating output itself, the draft model's job is to guess several tokens ahead cheaply, so the target model can check all of those guesses in one pass rather than generating each token through its own, slower forward pass.

What the draft model actually does

At each step, the draft model generates a short run of candidate tokens on its own, using whatever context has been accepted so far. Because it's small, it's fast enough to do this several times over with little cost. Those candidates are then handed to the target model, which scores all of them in a single forward pass: it checks, token by token, whether its own probability distribution would have picked the same token. Matches are accepted; the first mismatch is discarded, the target's own token is used instead, and drafting resumes from there. The draft model never gets a vote in what the final output is, it only proposes.

Why a small model can do this job

A draft model doesn't need to match the target model's quality, it needs to be right often enough on the easy parts of a sequence to save work, and cheap enough that being wrong sometimes doesn't cost much. Predictable spans, common phrasing, boilerplate code, repeated structure, are exactly where a much smaller model tends to agree with a much larger one, so the pairing works because language has a lot of low-entropy stretches a lightweight model can predict correctly.

How draft models are chosen

The most common approach pairs a target model with a smaller model from the same family, for example a distilled or lower-parameter variant trained on similar data, since shared training tends to produce more agreement between the two. Some serving stacks skip a second model entirely: they draft from the target model's own early layers exiting early, or from an n-gram table built out of the current context, both of which avoid loading a second model's weights into memory at all. Which approach is best depends on how much unified memory is available to hold a second model alongside the target, and on how much draft-model quality actually pays off for the workload in question.

The tradeoff that decides whether it's worth it

Every draft token costs some compute to generate, whether or not the target model accepts it. A draft model that's too weak gets rejected often, and the wasted drafting work can erase the speedup or make requests slower than plain decoding. A draft model that's too close in size to the target saves less relative compute per accepted token, even if its acceptance rate is high. Picking a draft model is a balance between those two failure modes, and it's worth testing against your own prompt distribution rather than assuming a published pairing transfers directly. Running a target and draft model concurrently also means budgeting unified memory for both, worth checking against what fits in memory before assuming a given pairing is workable on a single node.

Related pages

See single-stream throughput on a real Spark.

Measured tokens per second by concurrency, with the bandwidth arithmetic behind them.

Read the benchmarks