←Home KnowML
Home / Attention Lab

Attention Lab

One four-word sentence, run through seven different attention mechanisms. Every matrix comes from one seeded example β€” step through and watch the actual numbers, not a diagram of them.

Interactive Step through at your own pace Β· prerequisites: Attention & Transformers (08)
How to use this page β–Ύ

Pick a tab β€” MHA first if you're new to this, since every other tab reuses its vocabulary and its epilogue (concat, then one output projection). Every number below is computed by real matrix arithmetic in your browser from one seeded example (four tokens: "The cat sat down"), not hand-typed β€” the FlashAttention tab even checks its own answer against the plain-softmax result and shows you the difference.

β†’What this page simplifies, on purpose

  • No RoPE rotation is animated. Positions still matter β€” the causal mask depends on them β€” but the actual sinusoidal rotation of Q/K is left to page 08, so the numbers here stay focused on attention itself.
  • MLA's decoupled rotary key is described, not computed. The compression trick that saves memory is fully accurate; the small extra cached key component is mentioned in its own step rather than added to the arithmetic.
  • Batch is fixed at 1. Every tensor shown is for one sequence. Add a batch dimension and every shape below just gains an outer B Γ— β€” nothing about the math inside changes.
  • Every dimension is tiny. Real models use d_model in the thousands and dozens of heads. The mechanisms are identical at that scale β€” only the grids would no longer fit on a screen.