What is Temperature (LLM)?
Definition
Temperature is a sampling parameter that controls how sharply or how flatly a language model weighs its options when choosing the next token. Low values concentrate on the most likely tokens and give more consistent output; high values spread probability out and give more varied, less predictable output. Even a temperature of zero does not guarantee identical answers in every environment.
Also known as: sampling temperature, LLM temperature, top-p, nucleus sampling, top-k sampling

Reshaping the probability distribution
Temperature acts at every step of inference. For each token in its vocabulary the model produces a raw score, a logit. A softmax function turns those scores into probabilities, and the next token is drawn from that distribution. Temperature is the number every logit is divided by before the softmax. Values below 1 stretch the gaps between scores and make the favorite even more dominant; values above 1 squeeze the gaps and give unlikely tokens a better chance.
Take three candidate tokens with logits of 2.0, 1.0 and 0.5. Here is how their probabilities move:
| Candidate | T = 0.5 | T = 1.0 | T = 1.5 |
|---|---|---|---|
| A (logit 2.0) | 84.4% | 62.9% | 53.2% |
| B (logit 1.0) | 11.4% | 23.1% | 27.3% |
| C (logit 0.5) | 4.2% | 14.0% | 19.6% |
As temperature approaches zero, the distribution collapses onto a single token. That limit is greedy decoding: always picking the most probable option.
Top-p and top-k: trimming the tail
Temperature changes the shape of the distribution. Top-p and top-k limit which part of it the model may sample from.
- Top-k keeps only the k most likely tokens.
- Top-p, or nucleus sampling, sorts tokens by probability and keeps adding them until their combined probability reaches p. In the table above at T = 1.0, a top-p of 0.8 keeps A and B (62.9% + 23.1% = 86%), drops C, and renormalizes the remaining two.
OpenAI's API reference recommends adjusting temperature or top_p, but not both. A request for a task that needs consistency, such as classification, might look like this:
{
"model": "<model-name>",
"input": "Classify this review as positive, negative or neutral: ...",
"temperature": 0.2
}Same name, different ranges
The parameter shares a name across providers, but its range and even its availability vary. In OpenAI's Responses API, temperature runs from 0 to 2. In Anthropic's Messages API it runs from 0 to 1 with a default of 1.0, and Anthropic's API documentation notes that models released after Claude Opus 4.6 do not support setting it at all, accepting only 1.0 for backward compatibility. A "temperature: 0.7" carried from one provider to another therefore does not produce the same behavior, and sampling settings need retesting whenever you switch models.
Why zero is not a determinism switch
A common assumption is that temperature 0 makes a model return the same answer every time. Anthropic's documentation states the opposite plainly: even at 0.0, results will not be fully deterministic. Practical causes include floating-point operations completing in different orders on parallel hardware, requests being batched differently, and model updates on the provider's side. When two tokens are nearly tied, a tiny numerical difference can send the rest of the answer down another path.
If you need reproducibility, do not rely on temperature alone. Log the model version, the prompt and every parameter, validate output against a schema, and base important decisions on more than one run.
What it means for measuring AI visibility
Sampling randomness also matters when you try to measure how visible a brand is in AI answers. Asking an assistant the same question twice and getting two answers with different sources and different brand names is normal. A single screenshot is not reliable evidence either way. Metrics such as AI share of voice only become meaningful when the same prompts are run many times and results are reported as rates. Higher temperatures also make it more likely that a model wanders into improbable details, a risk discussed under hallucination.

