Sampling Lab
A model does not pick a word. It produces one row of numbers, and everything that decides which token you see happens afterwards. Step through that afterwards — softmax, temperature, top-k, top-p, and the draw itself — on one seeded example you can follow end to end.
How to use this page ▾
Start with Softmax — every later tab operates on the distribution it builds. Two prompts run throughout: a focused one ("The cat sat on the ___"), where the model is nearly certain, and an open one ("On Saturday I went down to the ___"), where a dozen continuations are about equally good. That pairing is the point: the Top-p tab shows a case where top-k and top-p choose genuinely different candidate sets, in opposite directions on the two rows, which is the whole argument for nucleus sampling. The Softmax tab also checks its own arithmetic — that the row sums to 1, that subtracting the max changes nothing, and that skipping it returns NaN.
→What this page simplifies, on purpose
- The vocabulary is 12 tokens, not 50,000. Every operation here — the softmax, the sort, the cumulative sum, the renormalisation — is exactly what runs over a real vocabulary; only the row width changes. What a real vocabulary adds is a very long, very flat tail, which is precisely what top-k and top-p exist to cut off, so the arguments on this page get stronger at scale, not weaker.
- The words are labels; the numbers are computed. The hidden states and the unembedding matrix are drawn from one seeded random stream, so the logits are whatever the arithmetic says. The candidate words are then attached in plausibility order — highest logit to the most plausible continuation — and displayed alphabetically so no row on the page arrives pre-sorted. Nothing numeric is hand-chosen.
- Two hidden states were selected, not invented. Each context picks its hidden state from a pool of 48 candidates drawn from the same seeded stream: the one landing closest to 85% top-1 confidence for the focused context, and the flattest of all 48 for the open one. That selection is the only editorial choice on the page — everything downstream of it is plain arithmetic.
- One position, one step. Real generation repeats all of this once per token, feeding each drawn token back in, which is where repetition penalties, beam search and speculative decoding live. This page stops at the first draw.
- Only the three settings everyone actually ships are covered. Temperature, top-k and top-p. Min-p, typical sampling, epsilon/eta sampling, Mirostat and contrastive search are all variations on the same two moves shown here: reshape the row, then truncate it before drawing.
- Sampling is done with a seeded PRNG.
mulberry32is not a cryptographic generator and a real server's RNG is not either — but "uniform on [0,1)" is all the inverse-CDF draw needs, and seeding it is what makes the counts on the last tab identical on every reload.