Neural Network Fundamentals
The chain rule, applied a few million times through a computational graph, is the entire training algorithm underneath every model on this site. The choices layered on top of it (activations, init, normalization, residuals) are the difference between a network that trains and one that does not.
A neural network is alternating linear transformations and non-linear activations. Backpropagation trains it by running the chain rule backward through the computational graph exactly once. Every weight gets its gradient in a single backward pass, no matter how deep it sits.
Everything else on this page is engineering built on that one algorithm: which non-linearity to use (ReLU/GELU/SiLU), how to initialize weights so signals neither vanish nor blow up before training starts (Xavier/He), how to keep activations well-scaled mid-training (BatchNorm/LayerNorm/RMSNorm), and how to let gradients skip past dozens of layers undamaged (residual connections).
One sentence for an interview: backprop is the chain rule, applied systematically to a computational graph, computing every parameter's gradient in one backward traversal.
01Intuition
Backpropagation answers one question, asked for every weight, at every step of training: how much is this specific weight to blame for the error?
A forward pass is a long chain of simple, differentiable steps. Multiply by a weight matrix, add a bias, squash through a non-linearity, repeat. Feed in an input and every layer transforms it a little, until you get a prediction.
Compare that prediction to the true answer and you get one number: the loss. "You were this wrong."
Computing that number was never the hard part. The hard part is credit assignment. For a weight sitting three layers deep, which direction should it move, and by how much, to make that one number smaller next time?
Brute force is hopeless: you would nudge one weight, rerun the entire forward pass, see what changed, and repeat for every weight, which does not scale to a million parameters.
A forward pass is a computational graph: a directed sequence of differentiable operations, each node's output feeding the next node's input.
The chain rule says the derivative of the loss with respect to any node is the product of local derivatives along the path from that node to the loss.
Here is the part that matters. Many paths share sub-paths. A weight in layer 1 affects the loss only through everything downstream of it, and that downstream work is identical for every weight in the layer. Compute it once, cache it, reuse it.
So one backward traversal gets you every weight's gradient. That reuse is the whole reason backprop is efficient and avoids an exponential cost.
Concretely, a 2-layer network is $x \to$ linear + activation $\to h \to$ linear + activation $\to \hat y \to$ loss $L$.
To find $\partial L / \partial W^{[1]}$, the gradient for the first layer's weights, you do not analyze layer 1 in isolation. You take the gradient signal that already arrived at the hidden layer, computed a moment ago while working out layer 2, and multiply by one more local derivative.
Gradients flow backward like water through a pipe network, picking up a multiplicative factor at every junction. This is also why they can shrink toward zero or blow up over many junctions. More on that in Tradeoffs.
A single neuron can't compute XOR, the function that is 1 only when exactly one input is 1, but two hidden neurons can. The weights below are hand-set to show that it works, and training would find something equivalent.
Try it Compute XOR with a two-neuron hidden layer and hand-set weights
import numpy as np
def relu(z): return np.maximum(0, z)
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
W1 = np.array([[1.0, 1.0], [1.0, 1.0]]) # both hidden units add inputs
b1 = np.array([0.0, -1.0]) # the second fires only above 1
W2 = np.array([1.0, -2.0]) # output = h1 - 2 * h2
out = relu(X @ W1 + b1) @ W2
for x, o in zip(X, out):
print(f"inputs {x.astype(int)} -> {o:.0f}")
0, 1, 1, 0 for the four inputs, which is XOR. The second hidden neuron switches on only when both inputs are 1, because their sum is then above 1. The output subtracts twice its value, which cancels the first neuron's 2 down to 0. A hidden layer is what lets a model make a decision that no single straight line can.02Timeline
Rosenblatt's perceptron (1958) learned only linearly separable patterns, and Minsky and Papert (1969) proved it could not learn XOR, which helped cause the first AI winter because nobody could train networks with hidden layers.
Rumelhart, Hinton and Williams (1986) popularized backpropagation, which computes the gradient for every weight in a multi-layer network with one forward and one backward pass and made hidden layers trainable.
Since then CNNs (LeCun, late 1980s–1998), RNNs/LSTMs and Transformers (2017) have changed the layers while the training algorithm stayed backprop, the chain rule and gradient descent.
03Architecture: forward pass, then backward pass through the same graph
A multilayer perceptron (MLP) is nothing more than this pattern, stacked: linear transformation → non-linear activation, repeated.
The universal approximation theorem (Cybenko, 1989; Hornik, 1991) says a single hidden layer, given enough neurons and a non-linear activation, can approximate any continuous function on a compact input domain to arbitrary precision.
The theorem is an existence proof and gives no training recipe. It says nothing about:
- How many neurons "enough" is. It can be exponentially many.
- Whether gradient descent will ever find those weights.
In practice depth matters, even though the theorem only requires width. A deep network can represent the same function with far fewer total neurons than a shallow one. Each layer composes on features the previous layer already built (edges → textures → parts → objects) where a shallow network would represent the whole function flat in one layer.
Real networks almost never stack raw linear+activation layers indefinitely. Past a modest depth, plain stacking becomes very hard to optimize (see Why, next).
The fix that made 50-, 100-, even 1000-layer networks trainable is the residual block: normalize, transform, and add the result back onto a running residual stream and does not overwrite it. This is the same residual stream that runs through every Transformer block running through every Transformer block (08).
Take shapes and cost concretely. A linear layer maps (batch, in_features) → (batch, out_features) via a weight matrix of shape (in_features, out_features) plus a bias. That is in_features × out_features + out_features parameters, regardless of batch size.
A useful rule of thumb from the scaling-law literature: total training compute is roughly $6 \times N \times D$, where $N$ is parameter count and $D$ is training tokens seen. It breaks down as:
- Forward pass: about $2N$ FLOPs, one multiply and one add per parameter.
- Backward pass: roughly twice the forward pass again, since it propagates gradients with respect to both the weights and the activations.
Two separate memory costs follow from that. Parameters set your floor for weights. Activations, which you must keep around for the backward pass, set a second cost that scales with batch size and depth, and it is often the larger of the two. This is why gradient accumulation and activation checkpointing exist (more in section 07 and in 23).
04The equations
Backprop through a 2-layer MLP
That is the forward pass. The backward pass computes a gradient $\delta$ at each layer's pre-activation, propagates it back through one linear map, and stops:
- δ the "error signal" arriving at a layer's pre-activation — everything the loss cares about that flows through this point, collapsed into one vector.
- f', g' is the local derivative of that layer's own activation function, evaluated at the value it output on the forward pass. This is why the forward pass's intermediate values have to be cached — the backward pass needs them.
- (W^[2])ᵀδ^[2] how the error signal at layer 2 propagates back through layer 2's weights to become "blame" felt by layer 1's output — this is the literal mechanism of credit assignment.
- δ(a)ᵀ an outer product: every weight $W_{ij}$ gets the product of "how wrong was output $i$" and "how active was input $j$" — a weight only gets blamed in proportion to how much it was actually used.
Derivation Deriving backpropagation from the chain rule alone
Those four equations were stated, not earned. Here is where each one comes from. Nothing below is more than the chain rule applied carefully, which is the point: backprop is not a separate algorithm you have to memorise.
-
$$\delta^{[l]} \;\equiv\; \frac{\partial L}{\partial z^{[l]}}$$Why define this at all: we want $\partial L / \partial W^{[l]}$, not $\partial L / \partial z^{[l]}$. But every parameter in layer $l$ influences the loss only by changing $z^{[l]}$. So if we know how sensitive $L$ is to $z^{[l]}$, one more short step gives us every gradient in that layer. $\delta$ is the quantity worth computing.
-
$$\delta^{[2]} = \frac{\partial L}{\partial \hat y} \odot g'(z^{[2]})$$Why the Hadamard product: the output activation acts elementwise, so $\hat y_i$ depends on $z^{[2]}_i$ and on no other component. The chain rule for a single index reads $\partial L/\partial z_i = (\partial L / \partial \hat y_i)\, g'(z_i)$. Stack those independent scalar equations and you get an elementwise product, not a matrix multiply.
-
$$\frac{\partial L}{\partial z^{[1]}_j} = \sum_i \frac{\partial L}{\partial z^{[2]}_i}\cdot\frac{\partial z^{[2]}_i}{\partial z^{[1]}_j}$$Why a sum: this is the multivariable chain rule. Perturbing $z^{[1]}_j$ changes $a^{[1]}_j$, which feeds into every unit of the next layer. The blame arriving at one hidden unit is the total collected over all the paths leaving it.
-
$$\frac{\partial z^{[2]}_i}{\partial z^{[1]}_j} = W^{[2]}_{ij}\, f'(z^{[1]}_j)$$Why: expand the single path. $z^{[2]}_i = \sum_j W^{[2]}_{ij} a^{[1]}_j + b^{[2]}_i$ and $a^{[1]}_j = f(z^{[1]}_j)$. Differentiating the composition gives the weight on that edge times the local activation slope.
-
$$\begin{aligned}\frac{\partial L}{\partial z^{[1]}_j} &= f'(z^{[1]}_j)\sum_i W^{[2]}_{ij}\,\delta^{[2]}_i \\[6pt] \delta^{[1]} &= \left(\left(W^{[2]}\right)^{\top}\delta^{[2]}\right)\odot f'(z^{[1]})\end{aligned}$$Why the transpose appears, concretely: in that sum, $i$ is the row index of $W^{[2]}$ and $j$ is fixed as the column. Summing down a column of $W$ is the same as summing across a row of $W^{\top}$. The forward pass used row $i$ to build output $i$ from all inputs. The backward pass reads column $j$ to collect what input $j$ is answerable for. Same matrix, read the other way.
-
$$\frac{\partial L}{\partial W^{[2]}_{ij}} = \delta^{[2]}_i\, a^{[1]}_j \;\;\Longrightarrow\;\; \frac{\partial L}{\partial W^{[2]}} = \delta^{[2]}\left(a^{[1]}\right)^{\top}$$Why the outer product: $W^{[2]}_{ij}$ appears in exactly one place in the whole forward pass, the expression for $z^{[2]}_i$, where it multiplies $a^{[1]}_j$. So its derivative is that single term, with no sum. Every $(i,j)$ pairing of the two vectors is an outer product.
-
$$\frac{\partial L}{\partial b^{[2]}} = \delta^{[2]}$$Why it is just $\delta$: $\partial z^{[2]}_i / \partial b^{[2]}_i = 1$. The bias enters additively, so the error signal passes through it unchanged.
Normalization
BatchNorm computes statistics per feature, across the batch:
LayerNorm uses the same formula, but $\mu$ and $\sigma^2$ range over each example's own features and never touch the batch dimension. RMSNorm goes one step further and drops mean-centering entirely, normalizing only by root-mean-square magnitude:
No batch statistics, no running mean, and in the original formulation no bias term or mean subtraction. Cheaper to compute than LayerNorm, and empirically about as effective. That combination is why it is the default in most modern LLMs.
05Why these specific choices
Why ReLU beat sigmoid/tanh. Sigmoid and tanh squash their input into a bounded range. For large positive or negative inputs the function is nearly flat, so its derivative is close to zero.
Chain a dozen of those together. The gradient reaching early layers is a product of a dozen near-zero numbers. It vanishes, and early layers stop learning almost entirely.
ReLU, $f(z)=\max(0,z)$, has a derivative of exactly 1 for any positive input, so it does not saturate or shrink gradients, and they pass through unchanged wherever the unit is active. This is why it unlocked deeper networks.
It has a cost: for negative inputs ReLU's derivative is exactly 0, so a unit that ends up permanently negative on every training example receives no gradient and dies. Two patches keep ReLU's advantage while fixing that:
- LeakyReLU. A small non-zero slope for negative inputs.
- GELU / SiLU. Smooth, with non-zero gradient almost everywhere.
Why GELU/SiLU in modern transformers. GELU is $x\cdot\Phi(x)$ and SiLU (Swish) is $x\cdot\sigma(x)$. Both multiply the input by a smooth gate: close to 0 for very negative inputs, close to 1 for very positive ones.
Unlike ReLU, the transition is smooth, and the function dips slightly negative just below zero and does not clamp hard at zero. That smoothness gives better-behaved gradients at the scale transformers are trained at, and empirically these activations outperform ReLU in that regime. Hence BERT, GPT-family models, and most modern transformer feed-forward blocks default to GELU or a gated SiLU variant (SwiGLU).
Why residual connections and not just a deeper plain network. He et al. (2015) showed something that looks paradoxical: a plain 56-layer network had higher training error than an 18-layer one. That is worse error on the data it was trained on, and overfitting does not explain it.
A deeper plain network is strictly more expressive, since it could always reproduce the shallow one by setting extra layers to the identity. So this is an optimization failure and not a capacity failure: plain SGD struggles to find that identity mapping across many stacked non-linear layers.
A residual block computes $f(x)+x$ in place of $f(x)$, so the identity mapping becomes the default whenever $f$ contributes little. The network starts at identity and only learns a useful deviation from it. This is what unlocked networks past ~20 layers.
06Normalization and regularization: what breaks
What you normalize over (batch, features, or a group of channels) is not a stylistic detail. It is the answer to a very common interview trap question: why do transformers use LayerNorm/RMSNorm and not BatchNorm?
| BatchNorm | LayerNorm | RMSNorm | GroupNorm | |
|---|---|---|---|---|
| Normalizes over | one feature/channel, across the batch | all features, within one example | all features, within one example (no centering) | a group of channels, within one example |
| Depends on batch size | Yes — unstable/wrong at batch size 1 | No | No | No |
| Needs running stats at inference | Yes (running mean/var — train/eval mismatch risk) | No | No | No |
| Typical home | CNNs / image classifiers | Transformers, RNNs, sequence models | Modern LLMs (LLaMA-family and similar) | Detection/segmentation — small per-GPU batch sizes |
BatchNorm's statistics are computed across examples in the batch. So the normalization applied to one example depends on which other examples happen to be sitting beside it.
This is fine for image classifiers, where batch size is typically large and fixed. It breaks badly for sequence models, for three reasons:
- Batch size varies between steps.
- Sequences are padded to different lengths.
- Critically, autoregressive generation happens one token at a time, with an effective batch dimension that does not resemble training at all.
LayerNorm and RMSNorm normalize within a single example's own feature vector. The computation is identical whether batch size is 512 or 1, and identical between training and inference. That independence from batch size and sequence length is the entire reason sequence models standardized on per-example normalization.
Failure modes that show up
- Dead ReLUs. A unit whose pre-activation is negative for every input in the training set receives zero gradient forever. A large learning-rate spike or an unlucky initialization can push a unit there, and it never recovers. Symptom: a growing fraction of activations reading exactly zero. Fixes: lower learning rate, He initialization, LeakyReLU/GELU/SiLU in place of plain ReLU.
- Vanishing / exploding gradients in RNNs. An RNN reuses the same weight matrix at every timestep, so the gradient flowing back through $T$ timesteps is roughly that matrix raised to the $T$-th power. Dominant eigenvalue below 1 and the gradient shrinks toward zero over long sequences. Above 1 and it grows exponentially, which can produce a NaN loss within a few hundred steps.
- Gradient clipping handles the exploding side. It caps the global gradient norm at a threshold before the optimizer step, which limits update magnitude without changing its direction. It is cheap, standard (Pascanu et al., 2013), and still routinely used when training any recurrent or otherwise deep, unstable model.
- Exploding logits. Unnormalized final-layer logits that grow very large before the softmax push cross-entropy toward numerical overflow, and make the loss landscape near that point extremely sharp. Handled in practice with the log-sum-exp trick, and in some modern LLMs with an explicit logit soft-cap.
"A deeper network takes proportionally longer to train per step because backprop has more layers to work through."
The backward pass costs roughly a fixed multiple of the forward pass, about 2×, regardless of how that compute is distributed across layers. Depth alone does not change that ratio.
Depth costs memory. Every layer's activations from the forward pass have to be kept around until the backward pass needs them. So memory is usually the first thing that breaks as you add layers, before per-step compute overhead. This is why activation checkpointing and gradient accumulation exist.
RMSNorm paired with a SwiGLU-style gated feed-forward block has become the default recipe across most frontier open-weight LLM families (LLaMA, Mistral, Qwen, and others). It has largely displaced the original 2017 Transformer's LayerNorm + ReLU/GELU feed-forward design.
The training algorithm underneath all of it is still exactly the backprop described on this page, largely unchanged since 1986, while the architecture keeps evolving.
Regularization: dropout, early stopping, weight decay
Normalization stabilises optimisation, while Regularization deliberately makes training harder so the network cannot memorise its way to a low loss. Page 01 frames the tradeoff; here is what the three most common levers do to the network.
- Dropout. During training, zero out each unit independently with probability $p$, then rescale the survivors by $1/(1-p)$ so the expected activation is unchanged. At evaluation time nothing is dropped. The effect is that no unit can rely on any particular other unit being present, which stops co-adaptation, where a group of units only works as a committee and none is individually meaningful.
- Early stopping. Track validation loss every epoch, keep the weights from the best one, and stop once it has not improved for a set number of epochs. It costs nothing and it is the only technique here that needs no hyperparameter tuning beyond the patience value.
- Weight decay. It adds a penalty proportional to the squared magnitude of the weights, which pulls them toward zero unless the data pushes back. This is $L_2$ regularization, and in a Bayesian reading it is a Gaussian prior centred on zero.
Dropout and BatchNorm interact badly, and it is a common interview follow-up. Dropout changes the variance of activations between training and evaluation, while BatchNorm's running statistics were estimated under training-time dropout. The two disagree at inference.
This is one reason modern transformer blocks use LayerNorm or RMSNorm with little or no dropout in the residual stream, and it is why model.eval() matters: it switches both layers into inference behaviour, and forgetting it is a classic source of results that are quietly wrong.
A practical ordering is to always use early stopping, since it is free, then add weight decay, because it is one number and it usually helps, and then add dropout when the model still overfits after both, and expect to need less of it than older tutorials suggest: large models trained on large corpora are often regularised adequately by the data itself.
07Build this
Section 05 makes a claim that sounds wrong. A deep plain network can have higher training error than a shallow one, on the same data. The loss curve is more convincing than the sentence.
Fit a smooth one-dimensional curve with a plain MLP, then keep making it deeper until it stops learning. Everything is hand-written from the four backprop equations in section 04, with no autograd. The aim is to watch the gradient shrink on its way back, and then change that by hand.
- Data: a few hundred points from a smooth curve with a couple of bends. One input, one output, no noise needed.
- Write forward and backward in NumPy from the section 04 equations. A $\delta$ at each pre-activation, an outer product for each weight gradient. Depth and width are arguments. Use tanh, and check the gradients once against finite differences before trusting anything.
- Train at depth 4, then at depth 40 with the same width, same learning rate, same number of steps. Plot both training-loss curves on one axis, plus each network's fitted curve over the target. If depth 40 still trains, keep adding layers until it stops. Where that happens is yours to find.
- Print the gradient norm for every layer of the deep net, output back to input, at the first step and again later in training.
- Add the residual: each block computes
x + f(x)in place off(x). In the backward pass that is one extra term, the incoming gradient passed straight through. Add hand-written RMSNorm from section 04 ahead of each block. Retrain at the same depth. - Now ablate: remove the residual and keep the norm, then keep the residual and remove the norm, with the same depth, seed and steps, and plot all three curves together.
Where this runs in production
Two of the tradeoffs above compound into a real constraint. A 7B-parameter model's weights, gradients, and optimizer state already consume tens of gigabytes in FP32, before a single activation is stored.
Three levers, each attacking a different part of that:
- Mixed precision. Store most of the forward and backward computation in FP16 or BF16. Half the memory, and much higher throughput on GPU tensor cores. Keep a master copy of the weights in FP32 so small updates are not lost to rounding, and add loss scaling to stop small gradients underflowing to zero. Loss scaling matters more for FP16's narrow exponent range than for BF16's.
- Gradient accumulation. This handles the batch-size side. If the batch you want does not fit, run several smaller micro-batches forward and backward, sum their gradients, and take one optimizer step after enough have accumulated. Mathematically close to training at the larger batch size, paid for with extra passes and not extra memory.
- Gradient clipping. Keeps the mixed-precision, accumulated updates numerically stable.
Together those are the standard recipe that lets a handful of memory-constrained GPUs train a model that would otherwise need far more hardware. The rest of the toolbox (sharding, offloading, activation checkpointing) gets a full treatment in 23.
08Interview questions
BeginnerWalk me through backpropagation for a simple 2-layer network.
Run the forward pass, caching every intermediate value: $z^{[1]}=W^{[1]}x+b^{[1]}$, $a^{[1]}=f(z^{[1]})$, $z^{[2]}=W^{[2]}a^{[1]}+b^{[2]}$, $\hat y=g(z^{[2]})$, then the loss. Going backward, compute $\delta^{[2]}=\partial L/\partial \hat y \odot g'(z^{[2]})$, use it to get $\partial L/\partial W^{[2]} = \delta^{[2]}(a^{[1]})^\top$, then propagate one layer further with $\delta^{[1]} = (W^{[2]\top}\delta^{[2]})\odot f'(z^{[1]})$ to get $\partial L/\partial W^{[1]}$. Each layer's gradient reuses the layer above it's already-computed signal, multiplied by one local derivative.
BeginnerWhy did ReLU replace sigmoid/tanh as the default activation?
Sigmoid and tanh saturate — for large-magnitude inputs their derivative approaches zero. Chained across many layers, that near-zero derivative gets multiplied together and the gradient reaching early layers vanishes, so those layers barely learn. ReLU's derivative is exactly 1 for any positive input, so gradients pass through unshrunk wherever the unit is active. The tradeoff is dead units: a unit stuck permanently negative gets exactly zero gradient and stops updating, which is what LeakyReLU/GELU/SiLU are designed to fix.
IntermediateWhat's the difference between BatchNorm and LayerNorm, and why do transformers use LayerNorm (or RMSNorm) instead of BatchNorm?
BatchNorm normalizes one feature's values across all examples in the batch; LayerNorm normalizes all of one example's features against each other, never touching the batch dimension. BatchNorm's statistics therefore depend on batch composition and size, need running estimates carried into inference, and break down at small or variable batch sizes. Sequence models have variable-length, often padded sequences and generate one token at a time at inference — batch statistics are unreliable or simply don't make sense there. LayerNorm and RMSNorm normalize per example, so they're identical regardless of batch size or between training and inference, which is why they're the standard for transformers and RNNs.
IntermediateWhy do residual connections make very deep networks trainable?
He et al. showed that plain (non-residual) deep networks can have higher training error than shallower ones — an optimization failure, since a deeper network could always represent the shallow function by learning identity in the extra layers, it just struggles to find that solution via SGD. A residual block computes $f(x)+x$ instead of $f(x)$, so identity is the default whenever the learned branch contributes little, and gradients get a direct additive path backward through the residual stream instead of having to survive being multiplied through every intervening non-linear transformation.
IntermediateXavier/Glorot vs. He initialization — what's the actual difference, and why does it matter?
Both scale the initial random weights so that activation variance stays roughly constant from layer to layer instead of shrinking or blowing up as depth increases. Xavier/Glorot initialization assumes a roughly linear, symmetric activation like tanh/sigmoid and scales variance by $1/\text{fan}_{avg}$. He initialization accounts for ReLU zeroing out roughly half of all pre-activations, so it scales variance by $2/\text{fan}_{in}$ to compensate for that lost variance. Using the wrong one for your activation function reintroduces exactly the vanishing/exploding signal problem initialization is supposed to prevent, just starting at step zero instead of building up over training.
DeepYou're training a 7B-parameter model and keep running out of GPU memory. What levers do you pull, and why?
Mixed precision (FP16/BF16 compute with an FP32 master copy of weights) roughly halves memory and increases tensor-core throughput. Gradient accumulation lets you simulate a larger batch size than fits in memory by summing gradients over several forward/backward micro-batches before one optimizer step. Gradient clipping keeps the resulting updates numerically stable under mixed precision. If that's still not enough, the next lever is activation checkpointing (recompute activations during the backward pass instead of storing all of them) and optimizer/parameter sharding — both covered in depth in 23. None of these change what backprop computes, they change how much of it you have to hold in memory at once.
DeepYour RNN's training loss suddenly goes to NaN a few hundred steps in. What's happening, and what do you do?
Classic exploding gradient: an RNN reuses the same weight matrix at every timestep, so the backward gradient is roughly that matrix raised to the power of the sequence length. If its dominant eigenvalue is above 1, the gradient — and eventually the loss — grows exponentially with sequence length until it overflows to NaN. Fix: clip the global gradient norm to a fixed threshold before the optimizer step (this caps magnitude without changing direction), consider a lower learning rate, and if this keeps recurring, consider whether the architecture itself (e.g. gated recurrence, or dropping recurrence for attention) is fighting you rather than a tuning issue.
IntermediateWhat does dropout do differently at training time and at evaluation time?
During training it zeroes each unit independently with probability $p$ and rescales the survivors by $1/(1-p)$, so the expected activation is unchanged. At evaluation nothing is dropped and no rescaling happens, which is why the scaling has to be applied during training rather than after. The mechanism prevents co-adaptation: no unit can depend on a particular other unit being present, so the network cannot rely on a committee where no member is individually meaningful. Forgetting model.eval() leaves dropout active at inference, which produces results that are noisy and slightly wrong rather than obviously broken.
DeepWhy do modern transformer blocks use little or no dropout in the residual stream?
Partly because dropout and normalization interact badly. Dropout changes activation variance between training and evaluation, and BatchNorm's running statistics were estimated under training-time dropout, so the two disagree at inference. Transformers avoid that by using LayerNorm or RMSNorm, which are per-example and need no running statistics. The larger reason is that regularization need scales with how much the model can memorise relative to the data. A large model trained on a very large corpus sees few repeats, so the data itself does much of the regularising and heavy dropout mostly slows convergence.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.