Scientific & Structured AI
A molecule turned upside down is the same molecule. Architectures that already know that need orders of magnitude less data than architectures that have to learn it from examples, and that single fact shapes the whole field.
Science has the opposite resource profile to vision and language. Data is scarce and expensive. Structure is abundant and exact: rotations, translations, permutations, conservation laws. The whole field is a trade of one for the other.
The mechanism that makes the trade is equivariance. Build a symmetry into the architecture and the model never spends capacity or samples learning it. That is why a network with 32 training molecules can beat one with 32 molecules and augmentation.
What breaks is verification. A model that emits candidates nobody can cheaply check produces a list, not a discovery. Page 37 argues that case across industries; this page is the machinery underneath it.
01Why this is a different game
Look at the diagram before the prose. Every architecture on this page is a choice about how far right to sit.
Vision and language models are trained on billions of examples that cost nothing to collect. Scientific data does not work like that.
One density-functional-theory energy is CPU-hours. One crystal structure is a lab. One protein structure was months of crystallography. The dataset you would like does not exist and cannot be bought.
What you have instead is knowledge about the problem that a web-scale engineer does not have:
- Exact symmetries. Translate a molecule, rotate it, or renumber its atoms. The energy does not change. This holds exactly.
- Conservation laws. Forces are the negative gradient of an energy. Mass and momentum are conserved. A model can be built so these hold by construction.
- Cheap verifiers. A predicted structure can be scored against a held-out experiment. A forecast is scored by tomorrow. A proposed equation can be evaluated.
Inductive bias is the name for assumptions a model carries before it sees any data. Convolution assumes translation matters less than local pattern. Attention assumes any token may matter to any other.
Data and structure are substitutes. Every symmetry you fail to build in becomes something the model must learn, and learning it costs examples you do not have.
In NLP this trade barely matters, because examples are free and the symmetries are weak. In chemistry it decides the outcome, and that asymmetry is why scientific ML reads like a different discipline and not a different application.
02Invariance and equivariance, precisely
Two words that get swapped casually. The difference decides what your model is allowed to output.
Let $g$ be a transformation from some group: a rotation, a translation, a relabelling of atoms. A function $f$ is invariant if the transformation makes no difference at all, and equivariant if the output transforms the same way the input did.
Energy is invariant. Force is equivariant. Rotate the molecule and the force vector rotates with it.
The groups that matter here are small and concrete:
- E(3): rotations, translations and reflections of 3D space. SE(3) is the same without reflections, which matters for chirality.
- Permutation of identical atoms or of graph nodes. Renumbering the atoms cannot change the energy.
- Gauge and scale symmetries in physics problems, which is why a PDE residual holds at every point rather than only where you sampled.
The cheapest route to invariance is to throw the pose away. Feed the network pairwise distances instead of coordinates and rotation becomes structurally impossible to detect. Here is that, in nine lines.
Try it Rotate a point cloud and watch one network change its mind
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(0, 1, (5, 3)) # a 5-atom point cloud
c, s = np.cos(0.7), np.sin(0.7) # 0.7 rad about z, ~40 degrees
R = np.array([[c, -s, 0.], [s, c, 0.], [0., 0., 1.]])
Xr = X @ R.T # the same molecule, turned
W1 = rng.normal(0, .4, (15, 8)) # 15 = 5 atoms x 3 coordinates
V1 = rng.normal(0, .4, (10, 8)) # 10 = the pairwise distances
W2 = rng.normal(0, .4, 8) # shared output layer
def coord_mlp(P): # reads raw xyz, the naive way
return np.tanh(P.reshape(-1) @ W1) @ W2
def dist_mlp(P): # reads |ri - rj| instead
D = np.linalg.norm(P[:, None, :] - P[None, :, :], axis=-1)
return np.tanh(D[np.triu_indices(5, 1)] @ V1) @ W2
print("coords", tuple(X.shape), " distances", (10,))
print("coord_mlp %.4f -> %.4f" % (coord_mlp(X), coord_mlp(Xr)))
print("dist_mlp %.4f -> %.4f" % (dist_mlp(X), dist_mlp(Xr)))
print("dist_mlp shifted by %.2e" % abs(dist_mlp(X) - dist_mlp(Xr)))
0.00e+00. coord_mlp moves from 1.4112 to 1.1209 for a molecule that did not change. dist_mlp returns 2.0109 both times, to the last bit. No training happened here and none is needed. Invariance came from the input representation, which is what "built into the architecture" means in practice.That trick has a price. Distances are scalars, so a distance network can only ever produce scalars. Ask it for a force vector and it has no way to point.
Equivariant networks solve that by carrying geometric objects through the layers rather than stripping them. Features are typed by how they transform: scalars, vectors, and higher-order tensors, written as irreducible representations of the rotation group.
Tensor Field Networks introduced the construction for 3D point clouds. Layers combine features with spherical harmonics through tensor products, and the type arithmetic guarantees the output rotates correctly. The e3nn library is the standard implementation.
Augmentation shows the model rotated copies and hopes it generalises. The guarantee is statistical, holds only near the poses you sampled, and has to be re-earned by every new layer.
Built-in equivariance is an algebraic identity. It holds for every rotation, at initialisation, before a single gradient step, and it cannot degrade during training. One is evidence. The other is proof.
03What the symmetry buys, measured
The argument is quantitative, so it should be measured rather than asserted. Two seconds of CPU is enough.
Fix a target that depends only on geometry. Train three models on the same 32 molecules: one on raw coordinates, one on raw coordinates with rotation augmentation, one on pairwise distances. Test on fresh molecules in arbitrary orientations.
Try it The data-efficiency gap, on 32 training examples
import torch, torch.nn as nn
torch.manual_seed(0)
def rot(P): # a random 3D rotation per sample
Q, _ = torch.linalg.qr(torch.randn(len(P), 3, 3))
return P @ (Q * torch.sign(torch.det(Q))[:, None, None]).mT
D = lambda P: torch.cdist(P, P) # (n, 4, 4)
dist = lambda P: D(P).flatten(1) # (n, 16), invariant
flat = lambda P: P.reshape(len(P), -1) # (n, 12), not
y_of = lambda P: torch.exp(-D(P)).triu(1).sum((1, 2))[:, None]
Ptr = torch.randn(32, 4, 3) # only 32 training molecules
Pte = rot(torch.randn(600, 4, 3)) # fresh ones, arbitrary poses
ytr, yte = y_of(Ptr), y_of(Pte)
def fit(feat, aug):
net = nn.Sequential(nn.Linear(feat(Ptr).shape[1], 64), nn.Tanh(),
nn.Linear(64, 1))
opt = torch.optim.Adam(net.parameters(), lr=0.02)
for _ in range(1500):
loss = ((net(feat(rot(Ptr) if aug else Ptr)) - ytr) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
return ((net(feat(Pte)) - yte) ** 2).mean().sqrt().item()
print("train", tuple(Ptr.shape), "| std(y) %.3f" % yte.std())
print("coords test RMSE %.3f" % fit(flat, False))
print("coords + aug test RMSE %.3f" % fit(flat, True))
print("distances test RMSE %.3f" % fit(dist, False))
std(y) = 0.411, which is what predicting the mean would score. The coordinate model gets 0.741. It is worse than a constant. Augmentation rescues it to 0.349, barely better than the mean. The distance model gets 0.078, about nine times better than the coordinate model, from identical data and a near-identical network. The only change is what the first layer is allowed to see.That gap is the field's central selling point, and it shows up in real systems. The NequIP paper reports matching or beating existing interatomic potentials with up to three orders of magnitude fewer training data.
Note also what augmentation did in the run above. It helped a lot and it did not close the gap, at a fixed number of steps.
Equivariance is a constraint, and constraints cost throughput and engineering effort. Brehmer et al. studied the trade directly and found that augmentation can close the data-efficiency gap given enough epochs, while equivariant models still won at every compute budget they tested.
The strongest counterexample is AlphaFold 3. Its diffusion module operates on raw atom coordinates with no equivariant processing at all, and the authors say plainly that they found none was required. Enough data and enough scale can substitute for the constraint. Most scientific problems do not have enough of either.
04Molecules: graphs, point clouds and forces
A molecule is a graph with 3D coordinates attached. Which of those two things you use decides what you can predict.
The 2D route treats the molecule as a graph of atoms and bonds and runs message passing over it. Page 18 covers that machinery. It is permutation-equivariant by construction and it never sees geometry.
For many property-prediction tasks that is enough, and the baseline to beat is older than any of it. Extended-connectivity fingerprints hash local substructures into a bit vector, feed it to gradient boosting, and win more often than the literature admits.
The 3D route uses coordinates, and the architectures form a clean ladder:
- SchNet uses continuous-filter convolutions over interatomic distances. Invariant, scalars only.
- DimeNet adds bond angles, because distances alone cannot distinguish some geometries.
- EGNN keeps coordinates as a feature that is updated by relative-position vectors, giving E(n) equivariance with almost no machinery.
- NequIP and MACE go fully equivariant with tensor features. MACE raises the body order of each message, so one layer captures many-body interactions.
The largest application is the machine-learned interatomic potential. Density functional theory gives accurate energies at a cost that makes long molecular dynamics impossible. A learned potential fits DFT energies and forces, then runs simulations orders of magnitude faster.
Do not predict forces with a separate head. Predict the scalar energy $E_\theta(\mathbf{r})$, then define the forces as its negative gradient.
Two things follow immediately. The forces are equivariant, because the gradient of an invariant scalar is a vector. And they are conservative, so energy is conserved along a trajectory. A separate force head gives you neither, and the simulation drifts.
Where this goes wrong is extrapolation. A potential fitted near equilibrium meets configurations no training point resembled, and molecular dynamics will find them.
A low force error on a held-out benchmark is not evidence that a 100-nanosecond trajectory stays stable. Those are different claims, and only the second one matters for the simulation you wanted.
05Protein structure prediction, mechanism first
Page 37 covers what AlphaFold changed and what it did not. This section is about how it works.
A protein is a chain of amino acids that folds into a shape, and the shape determines function. Predicting that shape from sequence was open for fifty years.
AlphaFold 2 combines several priors that are all worth naming separately:
- Coevolution. The input is not one sequence but a multiple sequence alignment of evolutionary relatives. Residue pairs that mutate together are probably in contact. Evolution did the labelling.
- Two coupled representations. The Evoformer maintains an MSA representation and a residue-pair representation, and passes information between them repeatedly.
- Geometric consistency. Triangle multiplicative updates constrain edge $ij$ using edges $ik$ and $jk$. Distances must obey the triangle inequality, so the update enforces something true rather than merely learned.
- Equivariant structure. Each residue gets a local frame. Invariant Point Attention reasons about points in those frames, so the module is equivariant to global rotation and translation.
- Recycling. The whole network is run again on its own output, three times, which is cheap depth.
It is trained end to end against a frame-aligned point error, and it emits a per-residue confidence, pLDDT, alongside the coordinates.
Protein language models take a different bet. ESM-2 is a masked language model trained on sequences alone, and ESMFold predicts structure from a single sequence with no alignment. It is much faster and generally less accurate on hard targets, which tells you how much the MSA was contributing.
AlphaFold 3 replaced the structure module with a diffusion module over raw atom coordinates, dropped the explicit equivariance, and extended coverage to ligands, nucleic acids and complexes.
A high pLDDT reliably marks confidently folded regions and a low one reliably marks disorder. That makes it excellent for triage.
It is not calibrated as "probability this residue is within 1 Å". Teams that threshold on it as though it were, then propagate the surviving structures into a downstream pipeline, inherit an error rate nobody measured.
06Designing molecules and proteins
Prediction scores a structure somebody gave you. Design produces one. The symmetry argument returns, in a sharper form.
Page 12 has the diffusion machinery. The scientific version changes one thing: the denoiser is equivariant.
The constraint has a purpose: if the noise is isotropic and the denoising network is equivariant, the learned distribution over structures is invariant to rotation. The model never spends capacity on orientation, because every orientation is already the same sample. Equivariant Diffusion for Molecule Generation in 3D is the reference construction.
Protein design runs as a pipeline rather than one model:
- RFdiffusion generates a backbone, optionally conditioned on a binding site or a motif to scaffold.
- ProteinMPNN solves the inverse problem: given that backbone, which amino-acid sequence folds into it?
- AlphaFold is then run as an in-silico filter. Does the designed sequence fold back to the requested backbone?
- The wet lab expresses the survivors and measures whether they do anything.
Step three is the interesting one. It is a cheap verifier bolted onto a generator, which is the same pattern page 37 traces across every field that has worked.
The honest caveat is that the wet-lab hit rate is the only number that counts, it varies enormously by target class, and a generator's reported validity rate tells you nothing about it.
You have the core argument and its two biggest applications. Symmetry is a substitute for data, invariant features buy scalars, equivariant features buy vectors, and forces should be gradients of an energy.
The rest of the page is the other half of scientific ML, where the structure you exploit is not a symmetry but a differential equation.
07Physics-informed neural networks
If you know the governing equation, you can put it in the loss. It is elegant and useful, and frequently the wrong tool.
Represent the solution as a network $u_\theta(t, x)$. Because the network is differentiable, you can evaluate the differential operator on it with autodiff, exactly and at any point. The residual of the PDE becomes a loss term.
Three terms: fit the observations, respect the boundary condition $\mathcal{B}$, and drive the PDE residual $\mathcal{N}[u_\theta]$ to zero. The residual is evaluated at collocation points you choose, not at data points. There is no mesh.
The payoff is that physics constrains the solution where no data exists. That is easiest to see by removing the data and watching what happens.
Try it Extrapolate an ODE far outside the data, using only the equation
import torch, torch.nn as nn
torch.manual_seed(0)
net = lambda: nn.Sequential(nn.Linear(1, 32), nn.Tanh(),
nn.Linear(32, 32), nn.Tanh(),
nn.Linear(32, 1))
td = torch.linspace(0, 2, 12)[:, None] # data exists only on [0, 2]
ud = torch.sin(td) # truth: u' = cos t, u(0) = 0
tc = torch.linspace(0, 8, 200)[:, None].requires_grad_(True)
tt = torch.linspace(2, 8, 61)[:, None] # scored where no data exists
def train(physics):
f = net()
opt = torch.optim.Adam(f.parameters(), lr=0.01)
for _ in range(4000):
loss = ((f(td) - ud) ** 2).mean()
if physics: # residual of the ODE itself
u = f(tc)
du, = torch.autograd.grad(u, tc, torch.ones_like(u),
create_graph=True)
loss = loss + ((du - torch.cos(tc)) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
return (f(tt) - torch.sin(tt)).abs().max().item()
print("data", tuple(td.shape), "collocation", tuple(tc.shape))
print("data loss only max |error| on [2,8]: %.4f" % train(False))
print("+ physics residual : %.4f" % train(True))
1.6343 against 0.0005, a factor of three thousand. Both networks saw the same 12 data points, all of them in [0, 2]. The plain fit leaves the data and produces nonsense: an error of 1.63 on a function whose entire range is $[-1, 1]$. The physics-informed one tracks sin out to $t = 8$ from 200 collocation points where nobody told it the answer. The residual term, not the data, is carrying it.That demo is also a fair warning, because the ODE there is trivial. Real problems behave worse, and the failure modes are well documented:
- Competing loss terms. Data, boundary and residual terms have different scales and different gradients. The weights $\lambda$ are hyperparameters, and a badly set one lets the network satisfy the residual while ignoring the boundary.
- Trivial minimisers. Many residuals are minimised by $u = 0$. If the data and boundary terms are weak, that is where the optimiser goes.
- Spectral bias. Networks learn low frequencies first, so stiff and multiscale problems converge slowly or not at all.
- Non-convex optimisation replacing a linear solve. For a problem a finite-element solver handles, you have swapped an exact method for gradient descent.
On forward problems that classical solvers handle well, PINNs are usually slower and less accurate. A 2023 study asked this directly in its title and found the finite-element method competitive or better across its benchmarks.
The real niche is the inverse problem. An unknown coefficient in the PDE becomes another trainable parameter, and sparse noisy measurements plus the equation recover it. That is awkward for a classical solver and natural here. Choose a PINN when you have partial data and a missing constant, and use a classical solver when you have a well-posed problem and a mesh.
08Neural operators: learning the solver, not the solution
A PINN is trained for one problem instance. Change the initial condition and you retrain. Operators fix exactly that.
A neural operator learns a mapping between function spaces. Give it an initial condition or a coefficient field as a function, and it returns the solution as a function. Training costs a lot of solved instances. Inference is a forward pass.
Two designs dominate:
- Fourier Neural Operator. Each layer transforms to Fourier space, multiplies a learned complex weight per mode, truncates the high modes, and transforms back. The FFT makes it a global convolution, so information crosses the domain in one layer.
- DeepONet. A branch network encodes the input function sampled at fixed sensors. A trunk network encodes the query location. The output is their inner product, a structure that follows a universal approximation theorem for operators.
The property worth understanding is discretisation invariance. Because the FNO's weights live on Fourier modes rather than grid cells, the same trained model can be evaluated on a finer grid than it was trained on. It learned an operator, not an array.
The same idea at scale is what weather models are. GraphCast and GenCast learn the map from one atmospheric state to the next, and page 37 has the deployment evidence and the skill numbers.
The costs are real. You need a large corpus of solved instances, which means running the classical solver many times first. The surrogate has no error bar, and it degrades silently off distribution, which is the regime you most wanted it for.
Two ways to use a known equation. A PINN puts the residual in the loss for one instance. An operator learns the solution map for a whole family, amortising the cost across instances.
What is left is the part where the model's output is something a human or a compiler can check, and the loop that decides which experiment to run next.
09Symbolic regression, and outputs you can check
Everything above outputs numbers. This outputs an equation, which a scientist can read and an algebra system can verify.
Symbolic regression searches the space of mathematical expressions for one that fits the data. The search space is discrete, combinatorial and enormous. Three approaches to it:
- Genetic programming. Mutate and recombine expression trees, keep what fits.
PySRis the practical modern implementation. - Physics-motivated decomposition. AI Feynman tests for properties that shrink the problem, including dimensional analysis, symmetry and separability, then recurses on the smaller pieces.
- Sparse regression. SINDy fixes a library of candidate terms and selects a sparse combination. This turns search into a regression with an $\ell_1$ penalty, which is dramatically more tractable.
The output is not one equation but a Pareto front trading accuracy against complexity. You pick the point. That choice is a scientific judgement and the method does not make it for you.
Be clear about the limits. Symbolic regression scales badly past a handful of variables, and it finds an expression that fits your data which may not be a law of nature.
The same shape, in two other fields
Look at what symbolic regression, theorem proving and program synthesis have in common. A vast discrete search space, and a checker that costs almost nothing: evaluate the expression, compile the Lean proof, run the tests.
That combination is where current AI is strongest. It is the argument page 36 makes for mathematics and page 37 makes for industry. Where the checker is a wet-lab experiment instead, throughput on candidates stops being the bottleneck.
10The loop around the model
In science the model rarely gets a fixed dataset. It gets to ask for the next experiment, which changes what a good model is.
Bayesian optimisation is the standard machinery for optimising something expensive. Fit a probabilistic surrogate, usually a Gaussian process, then choose the next point by maximising an acquisition function over it.
- Expected improvement scores how much a point is likely to beat the current best.
- Upper confidence bound adds a tunable multiple of the predicted standard deviation to the mean.
- Thompson sampling draws one function from the posterior and takes its argmax, so uncertainty drives exploration directly.
Active learning is the same idea aimed at the model rather than the objective. Label the points the model is least sure about. For neural potentials the usual uncertainty estimate is the disagreement across a small ensemble.
Inverse design inverts the question. You know the property you want and need a structure that has it. Either search a generative model's latent space against a property predictor, or condition the generator on the property directly.
Every acquisition function assumes the surrogate's uncertainty means something. A Gaussian process on twenty points is reasonable. A deep ensemble of neural potentials is often confidently wrong in exactly the regions it has never seen.
When that happens the loop stops exploring and repeatedly samples a region it believes is understood. The symptom is an active-learning campaign whose error curve flattens early while coverage of the design space quietly stays tiny.
11Build this
Section 03 showed the gap at one training-set size. The interesting object is the whole curve, and it takes an afternoon on a laptop.
Build a five-atom system with a known analytic energy, fit it with and without the symmetry built in, and then use each learned energy as a force field. The second half is where the differences stop being academic.
- Define ground truth: a Lennard-Jones energy over all pairs of five atoms in 3D. Generate configurations by perturbing a cluster, and keep the analytic forces too. You will need them later.
- Build three feature maps. Flattened coordinates, flattened coordinates with a fresh random rotation each epoch, and the sorted vector of pairwise distances. Same network body for all three.
- Train each at sizes 8, 32, 128, 512 and 2048, holding steps and architecture fixed. Plot test RMSE against training size on log-log axes. This is the figure this page is built around.
- Now get forces by differentiating the predicted energy with respect to the input coordinates, and skip adding a second head. Compare direction against the analytic force using cosine similarity.
- Run velocity-Verlet molecular dynamics for a few thousand steps under each learned force field. Plot total energy against time.
12What breaks
Almost none of these are modelling failures. They are evaluation failures that a modelling story got attached to.
- The random split leaked. Molecular datasets are full of near-duplicate scaffolds, so a random split puts relatives on both sides. Re-split by scaffold or by date and the headline number drops, sometimes by half.
- The potential is accurate and the simulation explodes. Force error on the held-out set looks excellent. Molecular dynamics wanders into configurations the training set never contained and the trajectory diverges.
- Symmetry was built in, then broken by preprocessing. A canonical orientation, an unsorted feature vector, or a global coordinate frame quietly reintroduces pose dependence that the architecture had eliminated.
- The PINN found a trivial solution. The residual is satisfied, the boundary term had a weight three orders of magnitude too small, and the network is happily returning something close to zero.
- The neural operator was only ever tested at training resolution. The discretisation-invariance claim is the selling point and is the thing nobody checked.
- The generated molecules cannot be made. Validity, novelty and synthesisability are three different properties. Only the first is usually in the loss.
- A confidence score is treated as a probability. pLDDT, ensemble variance and softmax outputs all rank well and none of them are calibrated out of the box.
- The result is a list with no hit rate. A model proposed 380,000 candidate crystals. The number that matters is what fraction survived contact with a chemist, and page 37 has that story in full.
13Where you meet this in the wild
Same inductive-bias argument, four very different verifiers.
Millions of candidate compounds, a few hundred that can be tested, and a verifier that costs weeks. The model's job is to rank, and the honest metric is the wet-lab hit rate of the top slice.
2D fingerprints and graph models as the cheap filter, 3D structure for the survivors.
Not an end-to-end generator pointed at the clinic.
Decades of uniform reanalysis data and a free honest verifier, because tomorrow arrives. It is the most favourable setup in the field, and the one that reached operations.
Learned surrogates for the full forecast, run alongside the physics model.
Screening needs simulations that DFT cannot afford, so a learned potential replaces it inside a larger search. Stability of the resulting dynamics matters more than any benchmark error.
Equivariant potentials plus active learning on the configurations the simulation visits.
Thousands of near-identical CFD or structural runs during design iteration. A surrogate trained on those runs turns a nightly job into an interactive one, inside the design envelope.
Neural operators for the repeated sweep.
Not for the certification run, where an unbounded error is unacceptable.
14Interview questions
BeginnerWhat is the difference between invariance and equivariance, and why does it matter?
A function is invariant to a transformation if applying the transformation to the input leaves the output unchanged, and equivariant if the output undergoes the corresponding transformation. Rotating a molecule leaves its energy unchanged, so energy is a rotation-invariant quantity, while the force on each atom rotates along with the molecule, so forces are equivariant. The distinction decides what a model is capable of emitting. A network built on pairwise distances discards the orientation entirely, which makes it exactly invariant but means it can only ever produce scalars, since it has no representation of direction. To predict a vector quantity you need equivariant internal features that carry orientation through the layers. Getting this wrong shows up as a model that cannot express the target at all, rather than as a model that trains badly.
BeginnerWhy is data augmentation not equivalent to building a symmetry into the architecture?
Augmentation teaches the symmetry statistically from examples, so the model approximately satisfies it near the transformations that were sampled and the guarantee degrades away from them. Built-in equivariance is an algebraic property of the architecture, so it holds exactly for every element of the group, at initialisation, before any training, and it cannot be lost during optimisation. The practical consequences are data efficiency and reliability. The augmented model spends capacity and training samples relearning something already known, and it needs more epochs to approach the same accuracy. Published work comparing the two finds that augmentation can eventually close the data-efficiency gap given enough epochs, while equivariant models still win at a fixed compute budget. Augmentation also cannot give you correct equivariant outputs such as forces, only approximately invariant scalars.
IntermediateWhy should a machine-learned force field predict energy and differentiate it, rather than predicting forces directly?
Because two physical properties come for free. If the network outputs a rotation-invariant scalar energy and forces are defined as the negative gradient of that energy with respect to atomic positions, the forces are automatically equivariant, since the gradient of an invariant scalar is a vector that rotates with the input. They are also automatically conservative, so no work is created around a closed loop and total energy is conserved along a molecular dynamics trajectory. A separate force head enforces neither property. It may be approximately consistent on the training distribution, but it produces a non-conservative field, and the symptom is steady energy drift or an outright blow-up during long simulations. The gradient approach costs one extra backward pass per evaluation, which is cheap against the failure it prevents.
IntermediateWalk through how AlphaFold 2 gets from a sequence to coordinates.
The input is not a single sequence but a multiple sequence alignment of evolutionary relatives, plus any structural templates, because residue pairs that mutate together are likely to be in physical contact and that coevolution signal carries much of the prior. The Evoformer trunk maintains two coupled representations, one over the alignment and one over residue pairs, and passes information between them for many blocks. Triangle multiplicative updates constrain each pair edge using the other two edges of every triangle, imposing the triangle inequality that real distances obey. The structure module then gives each residue a local frame and uses Invariant Point Attention to reason about points in those frames, making it equivariant to global rotation and translation, and it outputs backbone frames and torsion angles. The network is recycled on its own output several times, trained end to end against a frame-aligned point error, and emits a per-residue pLDDT confidence.
IntermediateWhat is a neural operator, and how does it differ from a physics-informed neural network?
A physics-informed network represents the solution to one problem instance as a network of the coordinates, trained by penalising the residual of the governing equation at collocation points. Change the initial condition or a coefficient and you train again from scratch. A neural operator instead learns a mapping between function spaces, taking an input function such as an initial condition or a coefficient field and returning the solution function, so one trained model covers a whole family of instances and inference is a single forward pass. The Fourier Neural Operator learns weights on Fourier modes, which makes each layer a global convolution and gives discretisation invariance, since the model can be evaluated on a finer grid than it was trained on. DeepONet uses a branch network over the input function sampled at fixed sensors and a trunk network over the query point. The trade is that operators need a large corpus of already-solved instances, which means running the classical solver many times first.
DeepWhen would you not use a PINN, and what is it genuinely good at?
Avoid it for well-posed forward problems an established numerical method already handles, because you are replacing an exact linear solve with non-convex gradient descent, and published comparisons find finite elements competitive or better on standard benchmarks in both accuracy and time. The specific failure modes are several loss terms at different scales whose relative weights are hyperparameters, residuals trivially minimised by the zero function when the data and boundary terms are too weak, spectral bias that makes stiff or high-frequency problems converge slowly, and an optimisation landscape that gets harder as the physics gets harder. The genuine niche is the inverse problem. An unknown coefficient in the equation becomes another trainable parameter, and sparse noisy measurements combined with the residual recover it without a mesh and without an adjoint solver, which is awkward classically and natural here. It is also reasonable where meshing is itself the hard part.
DeepYour equivariant model reports a strong test score. What would you check before believing it?
First check the split. Random splits over molecular datasets put near-identical scaffolds on both sides and inflate the score, so re-evaluate under a scaffold or temporal split and expect the number to fall. Second, test the symmetry rather than trusting it: pass a rotated, translated and atom-permuted copy of the same input and confirm the prediction is unchanged to numerical precision, because a canonical frame, centring or an unsorted feature vector routinely breaks a guarantee the architecture provided. Third, check the metric matches the use. Low force error on held-out snapshots does not establish that a long molecular dynamics trajectory stays stable, so run the dynamics and look at energy conservation. Fourth, compare against an unfashionable baseline such as fingerprints plus gradient boosting. Finally, examine where the errors live, since aggregate error hides systematic failure on the chemically novel subset you cared about.
DeepAlphaFold 3 dropped explicit equivariance from its structure prediction. Does that undermine the case for inductive bias?
It bounds the case rather than undermining it. The AlphaFold 3 diffusion module operates directly on raw atom coordinates without rotational frames or equivariant processing, relying on random rotation and translation augmentation instead, and the authors state that they found no invariance or equivariance requirement in the architecture. That is a real data point. With enough training data, capacity and compute, a network can learn a symmetry well enough that enforcing it costs more in complexity and throughput than it returns, and scaling studies show augmentation closing the data-efficiency gap given sufficient epochs. The condition is the part that matters. Most scientific problems have neither a Protein Data Bank nor that compute budget, and in the low-data regime the constraint still buys orders of magnitude, as equivariant interatomic potentials show. The question is not whether the symmetry is true, which it is, but whether you can afford to learn it.
15Go deeper
Primary sources for each mechanism on this page, including the papers that report the limitations.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.