←Home KnowML
Home / Tokenizer Lab

Tokenizer Lab

Seven words, twelve merges, one vocabulary β€” built in front of you. Every pair count, every merge decision and every token on this page comes from running byte-pair encoding, not drawn as a diagram of it. By the last tab you will know precisely why a model that reads tokens cannot count the letters in "strawberry".

Interactive Step through at your own pace Β· prerequisites: NLP Evolution (09)
How to use this page β–Ύ

Start on Train BPE β€” the other three tabs all use the vocabulary it builds. There is no random seed anywhere on this page: given the corpus and the tie-break rule, byte-pair encoding is deterministic, so every count you see is reproducible with a pencil. The Train and Encode tabs each end with a self-check that recomputes their result by a second, independent route and prints whether the two agree.

β†’What this page simplifies, on purpose

  • The base vocabulary is characters, not bytes. Real tokenizers start from the 256 byte values, which makes out-of-vocabulary impossible β€” any text at all, in any script, is some sequence of bytes. Starting from the characters this corpus happens to contain leaves a visible hole: the Encode tab shows cranberry needing letters the vocabulary has no id for. That hole is exactly what byte-level BPE closes.
  • No word-boundary marker. Production tokenizers attach the preceding space to the token (" berry" and "berry" are different ids) or append an end-of-word symbol, so merges can never cross a word boundary and word-final pieces stay distinct. Here each word is merged independently with no marker, which keeps the chips readable at the cost of that distinction.
  • Ties are broken by first appearance. In a corpus this small several pairs routinely tie for most frequent. Different libraries break ties differently, and a different rule gives a different β€” equally valid β€” merge list. The Train tab flags every tie as it happens.
  • No special tokens, no normalization, no pre-tokenization regex. A real pipeline lowercases or not, splits punctuation and digits by an explicit regex, and reserves ids for <|endoftext|> and friends before the first merge is ever applied. All of that sits around the algorithm shown here, not inside it.
  • Twelve merges, not fifty thousand. The loop is identical at production scale; only the size of the counting problem changes. The final tab's curve is the same curve real tokenizer designers read, computed on a corpus small enough to check by hand.