How it works
At each step a model scores every possible next token. Temperature reshapes those odds before one is picked: near 0 the likeliest token almost always wins, while higher values give less likely tokens a real chance. Top-p (nucleus sampling) is a related setting that limits the choice to the most probable tokens that together reach a given share.
Use a low temperature for extraction, classification, code and anything that must be consistent, and a higher one for brainstorming or creative writing. Even at 0, repeated answers are not guaranteed to match word for word. Ranges differ between providers (some accept 0 to 1, others 0 to 2), and some reasoning models fix the value or ignore it.
Related terms
More in AI and LLMs
Basics