Home / Labs
The Labs
Five worked examples, each small enough to follow one operation at a time and check by hand as you go.
Why these exist
Reading "the attention weights are a softmax over scaled dot products" is not the same as watching a 4×4 matrix of scores become probabilities whose rows sum to 1.
01Pick a lab
Roughly in the order they become useful. Each stands alone.
Where the gradients come from
Backprop Lab
A 2-layer MLP forward then backward, one chain-rule link at a time — then every analytic gradient checked against a numerical one.
Open the lab β
Why models cannot spell
Tokenizer Lab
Build a BPE vocabulary from a tiny corpus, merge by merge, and see why a word split into four tokens has no letter-level structure left.
Open the lab β
The one everything else is built on
Attention Lab
MHA, KV cache, MQA, GQA, MLA, PagedAttention and FlashAttention on one worked example, with every shape labelled.
Open the lab β
From a distribution to a word
Sampling Lab
Logits to probabilities, then temperature, top-k and top-p on the same row — including where top-k and top-p disagree.
Open the lab β
What you lose when you shrink it
Quantization Lab
Quantize a weight matrix to INT8 and read the error cell by cell, then watch one outlier consume the whole numeric range.
Open the lab β
02How to use them
- Predict before you advance. Say what shape the next matrix will be. Getting it wrong teaches more than getting it right.
- Watch the shapes, not the values. The numbers are seeded and arbitrary. The shapes are the whole point.
- Read the self-check steps. They are where a lab proves its own arithmetic instead of asserting it.
You have got it when you can state the shape of every tensor between the input embedding and the attention output, for a given sequence length, model dimension and head count — without opening the lab.