What is temperature in LLM sampling?
Temperature is a number, typically between 0 and 2, that controls how much randomness a language model uses when choosing its next token. At each generation step, the model computes a probability for every possible next token; temperature reshapes that distribution before a token gets sampled from it. A low temperature sharpens the distribution toward the highest-probability tokens, making output more deterministic and repetitive. A high temperature flattens it, giving lower-probability tokens a real chance of being picked, making output more varied and less predictable.
What actually happens at the sampling step
The model's raw output for each step is a set of scores (logits) across its whole vocabulary, one per possible next token. Temperature divides those logits before they're converted to probabilities: dividing by a value below 1 exaggerates the gap between the top candidates, making the model's favorite token even more dominant. Dividing by a value above 1 compresses that gap, letting less likely tokens compete more evenly. At temperature 0, sampling becomes deterministic: the model always picks its single highest-probability token, known as greedy decoding.
What it feels like in practice
Low temperature output tends to read as consistent and safe, sticking close to the most conventional phrasing, and running the same prompt twice produces near-identical output. High temperature output reads as more varied and creative, but past a certain point it starts producing outputs that are less coherent or drift off-topic, since the model is regularly picking tokens it considers less likely. Neither end is universally "better"; the right value depends entirely on whether the task rewards consistency (extraction, code, factual answers) or variety (brainstorming, creative writing, generating multiple distinct options).
Top-p and other sampling controls
Temperature is usually paired with other sampling parameters rather than used alone. Top-p (nucleus sampling) restricts the candidate pool to the smallest set of tokens whose cumulative probability crosses a threshold, cutting off the long unlikely tail regardless of how temperature reshaped it. Top-k does something similar by capping the candidate pool to a fixed number of tokens. These interact rather than stack independently, so changing one changes how the others behave; that makes sampling parameters worth tuning together for a given task rather than in isolation.
A note on reliability
No published benchmark figures are cited here deliberately, since the effect of a given temperature value is model-specific and task-specific; GPUwerk hasn't run controlled measurements across models to publish exact numbers. Anyone deploying a model in production should test a few values against their own task and data rather than reuse a default from an unrelated model or vendor.
Where temperature is set
Temperature is a request-level parameter in most serving setups, not a property baked into the model weights, so it can be changed per call without reloading or retraining anything. It's typically passed alongside the prompt in the API request, whether that's a hosted API or a self-hosted inference server. Because it's cheap to change, many production systems set it dynamically per request type, low for a structured-output endpoint, higher for a creative-writing one, rather than fixing one value for the whole deployment.
Why this matters for self-hosted deployment
Hosted API providers sometimes cap or hide certain sampling parameters, or change their defaults without much notice. Running a model on infrastructure you control removes that constraint: every sampling parameter the serving framework exposes, temperature, top-p, top-k, repetition penalties, is available to set and change as needed, and stays fixed unless you change it. That predictability matters for anyone tuning output behavior carefully for a specific task rather than accepting a provider's default.