Temperature & sampling
Turn the temperature dial on a language model. Watch top-k and top-p reshape its choices.
Turn the temperature dial on a language model. Watch top-k and top-p reshape its choices.
Prompt
Chai tastes best with a plate of hot samosas?
Several good answers: a little randomness keeps replies fresh.
No filter is on: every token keeps its chance, even the silly ones.
Sample 100 times
Draw 100 tokens from the final distribution and compare the counts with what the probabilities predict.
Mini story writer
A tiny model that only looks at the previous word, using the same settings. Three runs, 14 tokens each. Underlined words were unlikely picks (chance below 15%).
Once upon a time, the tiger saw the monkey jumped happily and sang loudly. The river. Suddenly
Different words: 92%
Once upon a time, the monkey jumped happily ever. Then a big tiger ate a tiny banana.
Different words: 92%
Once upon a time, the tiger slept. Suddenly the tiger roared and slept. The tiger ate a
Different words: 67%
A language model writes one token (a word or piece of a word) at a time. For every candidate it outputs a score called a logit. Softmax turns scores into probabilities: raise e to each score (so everything is positive and bigger scores win by a lot), then divide by the total so they add up to 1.
Temperature divides the logits first. Low T makes the favourite even more likely; at T = 0 the model is greedy and always picks it. That is reliable for facts but repetitive: the same context always gives the same next word, so text can loop forever. High T flattens the odds, so surprising words get picked: more creative, until it turns into nonsense.
Top-k keeps only the k best tokens; top-p (nucleus) keeps the smallest group of top tokens whose chances add up to p. Both cut off the long tail of silly options, then renormalise what is left so it sums to 1 again.
Takeaway: these settings do not make a model smarter; they only change how it picks. Low T for answers that must be right, moderate T with top-p ≈ 0.9 for stories and ideas.
| token | z | z/T | e^(z/T) | p | p′ |
|---|---|---|---|---|---|
| samosas | 3.1 | 3.1 | 22.2 | 0.354 | 0.354 |
| pakoras | 2.8 | 2.8 | 16.4 | 0.262 | 0.262 |
| biscuits | 2.3 | 2.3 | 9.97 | 0.159 | 0.159 |
| jalebis | 1.6 | 1.6 | 4.95 | 0.0789 | 0.0789 |
| rusks | 1.2 | 1.2 | 3.32 | 0.0529 | 0.0529 |
| toast | 1 | 1 | 2.72 | 0.0433 | 0.0433 |
| idlis | 0.4 | 0.4 | 1.49 | 0.0238 | 0.0238 |
| noodles | 0.1 | 0.1 | 1.11 | 0.0176 | 0.0176 |
| ice-cream | −0.8 | −0.8 | 0.449 | 0.00716 | 0.00716 |
| homework | −2 | −2 | 0.135 | 0.00216 | 0.00216 |
| total | not needed | 62.8 | 1 | 1 | |
A language model writes one token (a word or piece of a word) at a time. For every candidate it outputs a score called a logit. Softmax turns scores into probabilities: raise e to each score (so everything is positive and bigger scores win by a lot), then divide by the total so they add up to 1.
Temperature divides the logits first. Low T makes the favourite even more likely; at T = 0 the model is greedy and always picks it. That is reliable for facts but repetitive: the same context always gives the same next word, so text can loop forever. High T flattens the odds, so surprising words get picked: more creative, until it turns into nonsense.
Top-k keeps only the k best tokens; top-p (nucleus) keeps the smallest group of top tokens whose chances add up to p. Both cut off the long tail of silly options, then renormalise what is left so it sums to 1 again.
Takeaway: these settings do not make a model smarter; they only change how it picks. Low T for answers that must be right, moderate T with top-p ≈ 0.9 for stories and ideas.
Things to try