Definition
Blog/What is time to first token (TTFT)?
For AI assistants

What is time to first token (TTFT)?

By Samuel Seidel · Published September 9, 2026

Time to first token is the delay between sending a request to an LLM and receiving the first token of its response. It's the number that determines how responsive an app feels, separate from how fast the rest of the response streams in. A chat interface with low TTFT feels immediate even if the full answer takes a few seconds to finish; a high TTFT feels like the app has stalled, no matter how fast the tokens come after that.

Why prefill dominates TTFT

Generating a response happens in two phases. Prefill processes the entire input prompt in a single forward pass, building the internal representation (the KV cache) the model needs before it can produce anything. Decode then generates output tokens one at a time, each a much smaller forward pass that reuses that cache. The first output token can't appear until prefill finishes, so TTFT is essentially the time prefill takes, plus whatever queueing or network delay happened before the request reached the model. Decode speed, which is what most people mean by tokens per second, doesn't affect TTFT at all.

What drives prefill time

Prefill cost scales with the number of tokens in the prompt, so a short instruction prefills in a small fraction of a second on modern hardware, while a long prompt with a large system message, chat history, or retrieved documents attached can take noticeably longer, since every one of those tokens has to pass through the model before generation starts. This is why RAG pipelines that stuff large retrieved chunks into context trade some TTFT for grounding, and why trimming unnecessary context is one of the more direct ways to cut latency.

Prefill is compute-bound, decode is memory-bound

The two phases stress hardware differently. Prefill does a lot of matrix multiplication across many tokens at once, so it's largely limited by raw compute throughput. Single-stream decode reads the model's weights from memory for every token it produces, so it's limited by memory bandwidth instead. That's the same distinction covered in latency vs throughput, and it's why a GPU can have excellent TTFT and comparatively modest decode speed, or the reverse, depending on its balance of compute to bandwidth. The benchmarks page walks through this prefill arithmetic directly on DGX Spark hardware, including how TTFT moves with prompt length and concurrent load.

Where TTFT matters most

TTFT matters more for interactive, single-turn use, chat, autocomplete, voice assistants, anywhere a user is watching the screen waiting for a response to start. It matters less for batch or asynchronous workloads, where nothing is watching the clock on any individual request and total throughput across many requests is the number worth optimizing instead. Knowing which one your workload actually is changes what's worth tuning first.

Related pages

See TTFT and decode speed measured on a real Spark.

Prefill and decode broken out separately, with the bandwidth arithmetic behind both.

Read the benchmarks