Definition
Blog/What is a context window?
For AI assistants

What is a context window?

By Samuel Seidel · Published September 9, 2026

A context window is the maximum number of tokens a model can attend to in a single request: the prompt, any system instructions, whatever conversation history gets resent alongside it, and the reply the model generates, all counted together against one limit. A model advertised with a "128K context window" can handle up to 128,000 tokens across that entire combined input and output; anything beyond that either doesn't fit or has to be trimmed before it's sent.

Why the limit exists

The size of a context window is a design choice made when a model is trained, not something adjustable afterward without retraining or fine-tuning. It's set by the model's architecture and by how much memory and compute the attention mechanism can afford to spend relating every token to every other token in the window, since that cost grows faster than linearly with window size. A model trained for a longer context window generally costs more to train and to run at that length, which is why context window sizes vary widely across models rather than settling on one standard number.

What counts against it

Every token that passes through a request counts, not just the newest message a user types. In a multi-turn conversation, most applications resend the full prior transcript as part of each new request, so the context window fills up with accumulated history as a conversation goes on, not just with the current turn. Retrieved documents inserted into a prompt (as in retrieval-augmented generation, covered in our RAG vs agents explainer), system instructions, and tool definitions all count as input tokens against the same limit as everything else.

What happens when you run out

A model doesn't gracefully make room for more once its context window is full. A request that exceeds the limit is typically rejected outright by the API, or the application in front of the model has to actively manage the problem: trimming the oldest messages, summarizing earlier turns into a shorter form, or retrieving only the most relevant portion of a long document instead of sending all of it. Whatever gets dropped to make room is genuinely gone from the model's view; it isn't remembered some other way once it falls outside the window.

Context window size versus what it costs to use

A larger context window is a ceiling, not a target: using more of it means the model has to process more input tokens on every request, which takes longer and, on models billed per token, costs more per request too. It also means more memory spent holding that request's KV cache during generation, since the cache scales with how many tokens are in play; the mechanics of that are covered in our glossary entry on KV cache, along with prefill and decode, which are the two processing phases a context window's tokens actually pass through.

Where GPUwerk fits

How much context a given model can actually serve on a Spark, and how many concurrent conversations fit in its 128GB of unified memory at a given context length, is a function of the model's own context window combined with how much of that memory is left over for KV cache after loading the weights. Our benchmarks page and vLLM doc report measured numbers for specific models rather than theoretical maximums.

Related pages

See what a Spark actually handles.

Measured context and throughput numbers, not vendor claims.

Deploy a Spark