Home / Attention Lab
Attention Lab
One four-word sentence, run through seven different attention mechanisms. Every matrix comes from one seeded example β step through and watch the actual numbers, not a diagram of them.
How to use this page βΎ
Pick a tab β MHA first if you're new to this, since every other tab reuses its vocabulary and its epilogue (concat, then one output projection). Every number below is computed by real matrix arithmetic in your browser from one seeded example (four tokens: "The cat sat down"), not hand-typed β the FlashAttention tab even checks its own answer against the plain-softmax result and shows you the difference.
βWhat this page simplifies, on purpose
- No RoPE rotation is animated. Positions still matter β the causal mask depends on them β but the actual sinusoidal rotation of Q/K is left to page 08, so the numbers here stay focused on attention itself.
- MLA's decoupled rotary key is described, not computed. The compression trick that saves memory is fully accurate; the small extra cached key component is mentioned in its own step rather than added to the arithmetic.
- Batch is fixed at 1. Every tensor shown is for one sequence. Add a batch dimension and every shape below just gains an outer
B Γβ nothing about the math inside changes. - Every dimension is tiny. Real models use
d_modelin the thousands and dozens of heads. The mechanisms are identical at that scale β only the grids would no longer fit on a screen.