AI Engineering · interactive

How an LLM picks the next token

The model does not return an answer. It scores every possible next token, turns the scores into probabilities, trims the list, then rolls a weighted die. Change the settings and watch the odds move.

prompt

What is the capital of Pakistan?

model output so far

The capital of Pakistan is

  1. Logits
  2. ÷ T
  3. Softmax
  4. Top-K
  5. Top-P
  6. Renormalise
  7. Sample

Candidates for the next token

illustrative numbers

after temperaturefinal chance (after trimming)

Sampling = drop a random point on this strip. Wider slice, more likely.

Logits are raw scores. They can be any number, so softmax turns them into chances that add up to 100%: P = exp(logit / T) / Σ exp(logit / T).

Temperature divides the scores before softmax. Below 1 the gaps grow, so the top token wins more often. Above 1 the gaps shrink, so Lahore and Karachi get real chances.

Top-K and Top-P cut the long tail. What is left is scaled back up to 100%, because the die has to land somewhere.

The picked token is added to the text, and the whole loop runs again for the next token.