←Home KnowML
Home / Backprop Lab

Backprop Lab

One tiny network — 3 inputs, 4 tanh hidden units, 2 outputs — and one training example. Step through the forward pass, then walk the chain rule backward one derivative at a time, then watch all 29 of those derivatives get re-measured numerically and compared against the analytic ones.

Interactive Step through at your own pace · prerequisites: Neural Network Fundamentals (04)
How to use this page ▾

Take the tabs in order — Forward first, since every gradient on the later tabs is a derivative of the loss it computes. Each backward step shows three things side by side: the local derivative of that one operation, the incoming gradient from the step above it, and their product. That is the chain rule made visible rather than asserted.

The tab that matters most is Gradient check. It throws the whole derivation away and re-measures every partial derivative the slow way — nudge one parameter by h, re-run the entire network, see how much the loss moved — then prints the largest disagreement between the two methods. Every number on this page, including that verdict, is computed by real arithmetic in your browser from one seeded example. Nothing is hand-typed.

→What this page simplifies, on purpose

  • One hidden layer, not many. The backward pass here is nine steps long. A deep network repeats steps 5–8 once per layer and changes nothing else — the gradient of the input, ∂L/∂x, is precisely the tensor the layer below would receive as its incoming gradient.
  • The loss is half sum-of-squares, not cross-entropy. Squared error keeps the first gradient equal to the residual, so the chain is easy to follow all the way down. Softmax with cross-entropy produces the same "gradient equals residual" shortcut one layer later; everything behind that point is identical to what is shown here.
  • Plain gradient descent, no optimiser. The descent tab subtracts η · ∇L and stops. Momentum, Adam's per-parameter scaling, weight decay and learning-rate schedules all change what is done with the gradients, not how the gradients are computed.
  • No regularisation, no dropout, no normalisation layers. Each would add its own local derivative to the chain. None would change the structure of the chain.
  • Batch is 2, and only on the last tab. Big enough to show where the batch axis gets summed away and where it survives, small enough that every row still fits on screen.
  • Every dimension is tiny. Real networks have millions of parameters and would never print their gradients. They are also never gradient-checked at full size — the check is run on a small case exactly like this one, and that is the honest reason this page can run it at all.