←Home KnowML
TL;DR

A handful of results here are among the most consequential things computing has done for science. A protein-structure model won a Nobel Prize. AI weather forecasting is running operationally at Europe's forecasting centre. A randomised trial of 105,934 women found AI-supported mammography catches more cancers with 44% less radiologist reading time.

Several other famous results are weaker than their headlines. Section 01 gives you a four-question filter, and every section afterwards applies it honestly, including where the answer is unflattering.

One pattern runs through all of it, and it is the most useful thing on this page: AI wins where the search space is enormous and verification is cheap. It stalls where there is nothing to check answers against.

01How to read a claim

Four questions take a minute to ask and discard most of what you will read about AI in science.

  1. Is it peer-reviewed, or a blog post? Both can be right, but only one has had someone hostile read it carefully.
  2. Was it deployed, or demonstrated? A model that beats a baseline on a held-out set and a model running in production on live inputs are different achievements, and the gap between them is where most projects die.
  3. Compared against what? The relevant baseline is the best existing method operated competently, and a weak version of it or no method at all would be too easy.
  4. Did anyone independent reproduce it? Reproduction is where several of the results in this page's own list came apart.
The number that should calibrate you

A 2026 review looked at 1,357 AI-enabled medical devices cleared by the FDA. Of those, 34 were linked to a registered prospective trial. Twelve had a peer-reviewed publication. Three had been evaluated against patient-centred outcomes such as mortality, morbidity or readmission.

Regulatory clearance is not evidence of clinical benefit. It means the device was shown to be substantially equivalent to something already on the market. Read "FDA-cleared" as "legally sellable" and do not read it as "proven to help patients".

02Structural biology: the strongest case there is

If you want one example of machine learning changing a scientific field outright, this is it.

A protein is a chain of amino acids that folds into a specific three-dimensional shape, and the shape determines what it does. Determining that shape experimentally, by X-ray crystallography or cryo-electron microscopy, historically took months to years per protein. Predicting it from sequence was a fifty-year open problem with a biennial competition, CASP, tracking how badly everyone was doing.

AlphaFold 2 ended that in 2020: at CASP14 it reached accuracy comparable to experimental structures for most targets. The consequences were unusually concrete:

  • Coverage went from thousands to everything. The AlphaFold Protein Structure Database now holds predicted structures for over 200 million protein sequences, almost every protein known to science, released openly.
  • Adoption was immediate and broad. More than two million researchers across around 190 countries have used it.
  • It won the Nobel Prize in Chemistry in 2024, shared by Demis Hassabis and John Jumper for AlphaFold and David Baker for computational protein design.
Why this problem was winnable

The ingredients explain it. Fifty years of experimentally determined structures in the Protein Data Bank gave a large, high-quality, physically grounded training set. CASP gave an adversarial, blind, independently scored benchmark that nobody could game. And the target was a single well-posed output: coordinates.

Labelled data, an honest verifier, one clear output. When a scientific problem has that shape, expect machine learning to make serious progress. When it does not, be sceptical of anyone promising the same.

What it did not do is worth stating plainly, because it is routinely oversold. AlphaFold predicts a structure, typically one conformation. Real proteins move between states, and function often depends on that motion.

Predictions for intrinsically disordered regions, large allosteric transitions and ligand binding remain much weaker. AlphaFold 3 extended coverage to complexes, nucleic acids and ligands, and drug discovery still needs dynamics and binding affinity that a static structure does not give you.

03Weather: the strongest case for deployment

The clearest example of AI displacing a physics simulation in genuine operational use, at a national-scale institution, with published skill scores.

Numerical weather prediction integrates the equations of atmospheric physics on a supercomputer. It is one of computing's great successes and it is expensive: a global forecast is hours of compute on a dedicated cluster drawing megawatts.

Learned models changed the economics, and then the accuracy.

  • GraphCast (Science, 2023) showed a graph neural network could beat the operational deterministic baseline on most variables and lead times, running in minutes.
  • GenCast (Nature, December 2024) went further and probabilistic. It outperformed Europe's leading ensemble system on 97.2% of 1,320 evaluation targets, producing a 15-day global forecast in about 8 minutes on a single TPU.
  • It is operational and is more than a demo. ECMWF put its own AI Forecasting System, AIFS, into operations on 25 February 2025, running alongside the physics-based system, with reported gains of up to 20% on measures including tropical cyclone tracks.

This one passes all four questions in section 01. Peer-reviewed, deployed at the institution whose forecasts feed European national weather services, measured against the best existing system, and independently run by that institution and not only by the model's authors.

The caveat concerns what these models are trained on. They learn from decades of reanalysis data, so their competence is anchored in the climate that produced that data. How they behave on unprecedented events is an active research question, and unprecedented events are the regime where forecasting matters most.

04Materials: where the headline outran the result

A useful case study in reading a claim, because the work is real and the framing was not.

In November 2023 DeepMind published GNoME in Nature, reporting 2.2 million predicted crystal structures, of which around 380,000 were predicted stable. The announcement described this as discovering millions of new materials and compared the haul to centuries of accumulated human knowledge.

Materials scientists then looked at the output. Anthony Cheetham and Ram Seshadri of UC Santa Barbara sampled the released structures and applied a three-part test: is each proposed compound credible, useful, and novel? They found no strikingly novel compounds in the sample, and characterised much of the set as combinatorial variations and orderings of known compositions.

"Discovery" is doing a lot of work in that sentence

Both things are true at once. GNoME predicts stability at a throughput no human workflow could match, and that is a real contribution to a real bottleneck. And a list of 380,000 plausible crystals is not 380,000 discoveries, because a discovery requires that somebody wanted the material, can make it, and finds it does something new.

The general lesson: when a result is a list of candidates, the interesting number is never the length of the list. It is the hit rate after somebody tries to use it.

The scope is also narrow, because the predictions are inorganic crystalline compounds. Polymers, glasses, metal-organic frameworks, composites and heterostructures, which is where a great deal of applied materials work happens, were not addressed.

05Chip design: real deployment, contested benchmark

A case where the company is almost certainly right about its own use and the external evidence is thin.

Placing blocks on a chip die, floorplanning, is a combinatorial optimisation problem that human engineers spend weeks on. DeepMind published a reinforcement-learning approach in Nature in 2021, later branded AlphaChip, and reports using it in the design of multiple generations of Google's TPUs.

Both halves of the story matter:

  • Nature investigated and upheld it. After criticism, the journal completed a post-publication review in 2024 and published an Addendum in September 2024 supporting the original work.
  • Independent verification remains limited. Critics including Rice University's Moshe Vardi have said plainly that the community has not been able to verify the claims, and competing benchmark studies reached less favourable conclusions.

The reasonable reading: the method is in production at the organisation that built it, which is meaningful evidence, and the published comparisons are not strong enough to tell you how it would perform in your flow. This situation is common and under-acknowledged for industrial AI results, where the deployment is proprietary and the benchmark is a proxy.

06Physical control: fusion plasma

The most technically striking category, and the one with the least commercial noise around it.

A tokamak confines plasma at over a hundred million degrees using magnetic fields, adjusted thousands of times a second. Designing those controllers has traditionally meant hand-built control laws for each plasma shape.

DeepMind and EPFL published work in Nature in 2022 in which a reinforcement-learning policy commanded the full set of control coils on the TCV tokamak in Lausanne. It held conventional elongated shapes and advanced configurations including negative triangularity and a "snowflake", and sustained two separate plasmas simultaneously in the vessel.

What makes it notable is not the physics score. It is that the policy was trained in simulation and then ran on the real machine. Crossing that sim-to-real gap on a device where a control failure damages hardware is one of the hardest things reinforcement learning has been asked to do. Page 21 covers why that transfer is so difficult.

07Medicine: the one that reached patients properly

Thousands of medical AI products exist. Very few have been tested the way a drug would be, and this is what it looks like when one is.

The MASAI trial in Sweden randomised 105,934 women to either AI-supported mammography screening or standard double reading by two radiologists. It is a proper randomised, controlled, population-based screening trial, and the results have been published across The Lancet family of journals.

MeasureAI-supportedStandard double reading
Interval cancer rate (per 1,000)1.551.76
Specificity80.5%73.8%
Screen-reading workload44% reduction in reading burden (interim safety analysis, 2023)

Read the interval cancer number carefully, because it is the one that matters most. An interval cancer is one that shows up between screening rounds, meaning screening missed it. Fewer interval cancers means the programme is catching disease earlier, which is the actual point of screening. Higher specificity means fewer women called back for a scare that turns out to be nothing.

Set this against the calibration figure from section 01. Out of 1,357 cleared AI medical devices, three had been evaluated on patient outcomes. MASAI is the exception that shows what the standard could be, and it is no evidence that the field generally meets it.

08Algorithms: AI improving the tools it runs on

A small category with an outsized intellectual payoff, and a direct callback to page 36.

  • AlphaTensor (Nature, 2022) found faster matrix multiplication algorithms by treating algorithm discovery as a game.
  • FunSearch (Nature, 2024) used a language model inside an evolutionary loop to improve a bound on the cap set problem, a genuine open question in combinatorics.
  • AlphaEvolve (2025) found a scheme multiplying two 4×4 matrices in 48 scalar multiplications, improving on the 49 that Strassen's algorithm had held since 1969.

Fifty-six years is a long time for a bound to stand in a problem this heavily studied. And the reason this category works is the same reason the mathematics results in page 36 work: a candidate algorithm can be checked mechanically and exactly. The model is free to be wrong almost all the time.

09The pattern worth taking away

If you remember one thing from this page, make it this. It predicts which AI projects succeed better than any list of applications.

Go back through the sections and sort them by how well they worked.

Every result on this page, placed by the two axes that predicted the outcome
CHEAP VERIFIER · AI EATS THIS COSTLY VERIFIER · HYPOTHESES ONLY cost of checking one answer → size of the search space → huge small AlphaFold scored against known structures weather forecasting graded by tomorrow AlphaEvolve checked by arithmetic Lean-verified proofs a compiler says yes/no Riemann, Hodge, BSD no verifier for the right idea GNoME · materials somebody must synthesise it drug candidates a multi-year trial tokamak control simulator, then real machine chip floorplanning scored by place-and-route mammography screening only an RCT settles it cheap, automatic, honest a lab, a trial, or an expert's opinion
Position within a quadrant is editorial; which quadrant something lands in is the argument. Reading left to right predicts the outcome better than anything about the models involved: everything on the green side shipped, everything on the red side produced a list somebody still has to test. The vertical axis matters separately — chip floorplanning has a cheap verifier but a narrower search space, so it is useful rather than transformative. If you are choosing what to work on, moving a problem leftward by building a cheap verifier for your field is often worth more than a better model.
Where it wins

Huge search space, cheap verifier. Protein structures scored against experimental data. Weather forecasts scored against what happened tomorrow. Matrix multiplication schemes checked by arithmetic. Lean proofs checked by a compiler. The model can be unreliable, because everything wrong gets filtered out for free.

→
Where it struggles

No verifier, or an expensive one. Is this crystal a useful material? Somebody has to synthesise it. Does this drug work? Run a trial for years. Is this the right definition for attacking Riemann? Nobody knows how to check that. Prediction throughput stops being the bottleneck.

→
What that tells you

When you evaluate any AI-for-science claim, find the verifier first. Ask what checks the model's output, how much that check costs, and how honest it is. If there isn't one, the result is a list of hypotheses, which has value but is not a discovery. Most disappointment in this field is a verifier problem wearing a modelling costume.

10What breaks

The recurring failure modes mostly concern evidence and not models.

  • Clearance mistaken for proof. "FDA-cleared" means sellable, and three out of 1,357 cleared devices had been checked against patient outcomes.
  • Candidate lists counted as discoveries. 380,000 predicted stable crystals is a starting point. The number that matters is how many were made and turned out to be useful.
  • The benchmark is not the job. Chip floorplanning scores on public benchmarks tell you little about a proprietary production flow, in either direction.
  • Distribution shift where it hurts most. Weather models learn from historical reanalysis, and the events people care about are the ones with the least precedent.
  • One structure is not one molecule. A single predicted conformation omits the dynamics that determine binding, which is most of what drug discovery needs.
  • The demo-to-deployment gap is where projects die. Weather succeeded partly because forecasting centres already had the ingestion, verification and operational discipline to absorb a new model, which most fields do not.
  • Credit gets compressed. Results built on decades of human technique get reported as the machine's achievement, which is both unfair and misleading about how the next one will happen.

11If one of these is what you want to work on

The point of this page. Each of these fields is reachable from material on this site, and here is the actual route.

Structural biology & drug discovery

Attention over sequences, equivariant architectures, and generative models for designing molecules as well as predicting them.

Start with 08, then 12 and 19.

Weather & climate

Graph neural networks over the globe, diffusion for probabilistic ensembles, and time-series evaluation done properly.

Start with 18, then 17 and 12.

Fusion, robotics & control

Reinforcement learning, simulation, and the sim-to-real transfer problem that decides whether any of it reaches hardware.

Start with 15, then 21 and 22.

Medical imaging

Convolutional and vision-transformer architectures, calibration, and the evaluation discipline that separates MASAI from the other 1,354.

Start with 05, then 06 and 25.

Algorithm & program discovery

Search with a verifier in the loop: reasoning models, reinforcement learning against a checker, and evolutionary outer loops.

Start with 15, then 11 and 36.

The systems underneath all of it

None of these results happened without someone making enormous models train and serve efficiently. It is the least glamorous and most transferable skill on the list.

Start with 23, then 28 and 31.

12Interview questions

BeginnerWhat did AlphaFold change, and what did it not?

It made predicting a protein's three-dimensional structure from its amino-acid sequence accurate enough to substitute for experiment in many cases, collapsing a task that took months of crystallography into minutes of compute. The AlphaFold database now covers over 200 million sequences openly, it has been used by more than two million researchers, and the work shared the 2024 Nobel Prize in Chemistry. What it did not do is solve protein behaviour. It typically predicts a single conformation, while real proteins move between states, and predictions remain much weaker for disordered regions, large allosteric changes and ligand binding. Drug discovery needs dynamics and binding affinity, which a static structure does not provide.

BeginnerWhy is "FDA-cleared" weak evidence that a medical AI tool helps patients?

Because clearance under the common pathway establishes substantial equivalence to a device already on the market, not clinical benefit. A 2026 review of 1,357 cleared AI-enabled devices found 34 linked to a registered prospective trial, 12 with a peer-reviewed publication, and only three evaluated against patient-centred outcomes such as mortality, morbidity or readmission. So clearance tells you a product may legally be sold, and says almost nothing about whether using it improves outcomes. The MASAI mammography trial, which randomised over 105,000 women and reported both a lower interval cancer rate and higher specificity, shows what the stronger standard looks like and how rarely it is met.

IntermediateWhy did AI weather forecasting succeed in operational deployment when many AI-for-science projects stall?

Because every prerequisite happened to be in place. There is a large, uniform, high-quality training corpus in decades of reanalysis data. There is a free and completely honest verifier: tomorrow arrives and you score the forecast against it. The baseline is well defined and rigorously tracked, so improvements are measurable rather than arguable. And forecasting centres already had the operational machinery to ingest, verify and serve a new model, which is usually the part that kills a project. GenCast beat the leading ensemble on 97.2% of 1,320 targets while running in about eight minutes on a single accelerator, and ECMWF put its own AI system into operations in February 2025. The scarce ingredient in most other fields is the cheap honest verifier, not the model.

IntermediateWhat was the substance of the criticism of the GNoME materials results?

That the headline conflated predicted candidates with discoveries. DeepMind reported 2.2 million predicted crystal structures with around 380,000 predicted stable, framed as discovering millions of new materials. Cheetham and Seshadri sampled the release and applied a three-part test of credibility, usefulness and novelty, reporting that they found no strikingly novel compounds in their sample and that much of the set consisted of combinatorial variants and orderings of known compositions. They also noted the scope was inorganic crystalline compounds, excluding polymers, glasses, metal-organic frameworks and composites. The defensible version of the result is that stability prediction at very high throughput is genuinely useful; the indefensible version is counting list length as discoveries.

IntermediateWhy is the fusion plasma control result significant beyond fusion?

Because it crossed the sim-to-real gap on a system where failure is expensive. A reinforcement-learning policy was trained against a tokamak simulator and then commanded the real control coils on the TCV device, holding conventional and advanced plasma configurations including negative triangularity and a snowflake, and sustaining two plasmas at once. Most reinforcement learning successes are in simulation or in games, where the environment and the deployment target are identical. Here the policy had to survive the mismatch between a model of the physics and the actual machine, under hard real-time constraints, on hardware that damage matters to. That transfer problem is the central obstacle to reinforcement learning in the physical world, which is why the result generalises well beyond the application.

DeepYou are asked to assess a new AI-for-science claim. What framework would you use?

Find the verifier first: what checks the model's output, how expensive is that check, and is it independent of the model's authors. If verification is cheap and honest, as with a forecast scored against tomorrow's weather or an algorithm checked by arithmetic, then unreliable generation is fine because errors are filtered for free, and the result is likely to hold up. If verification requires synthesis, a clinical trial or expert judgement, the model has produced a hypothesis list and the meaningful number is the downstream hit rate, which is usually unreported. Then apply the standard four checks: peer-reviewed rather than announced, deployed rather than demonstrated, compared against the best existing method competently operated, and independently reproduced. Finally ask what the claim would look like if it were false, and whether the evidence presented would distinguish those cases.

DeepWhich scientific problems should you expect machine learning to make progress on next?

The ones whose structure resembles the successes rather than the disappointments. Look for three things: a large body of consistent, well-measured historical data; a well-posed output the model can be asked for; and a verification signal that is cheap, fast and not gameable. Protein folding had all three, with the Protein Data Bank, coordinates as output, and CASP as a blind benchmark. Weather had all three, with reanalysis archives, gridded forecasts, and tomorrow as the judge. Conversely, expect slow progress where the verifier is a wet-lab experiment, a multi-year trial or a human expert's opinion, because throughput on hypotheses stops being the constraint. The corollary is that building a cheap verifier for a field is often a higher-leverage contribution than building a better model for it.

13Go deeper

Primary sources, including the critical ones. Read the critiques alongside the announcements.

📄
Paper
Highly accurate protein structure prediction with AlphaFold
Jumper et al., Nature 2021. The paper behind the 2024 Nobel Prize in Chemistry.
📝
Primary
Nobel Prize in Chemistry 2024 — press release
The committee's own account of why structure prediction and protein design were recognised together.
📄
Paper
Probabilistic weather forecasting with machine learning
The GenCast paper, Nature 2024 — where the 97.2%-of-1,320-targets figure comes from.
📝
Primary
ECMWF — AI forecasts become operational
The forecasting centre's own announcement, February 2025. Deployment evidence rather than benchmark evidence.
📄
Paper
Scaling deep learning for materials discovery
The GNoME paper, Nature 2023. Read it next to the critique below rather than on its own.
📝
Article
AI is dreaming up millions of new materials. Are they any good?
Nature news on the credibility, usefulness and novelty question. The counterweight to the headline.
📄
Paper
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave et al., Nature 2022. Trained in simulation, run on the real TCV tokamak.
📄
Paper
MASAI: interval cancer, sensitivity and specificity
The Lancet. A randomised controlled screening trial of 105,934 women — the standard other medical AI is rarely held to.
📄
Paper
Addendum: A graph placement methodology for fast chip design
Nature's September 2024 addendum following post-publication review of the AlphaChip work.
📄
Paper
AlphaEvolve: A coding agent for scientific and algorithmic discovery
The 4×4 matrix multiplication result, and the clearest illustration of search plus verifier — arXiv:2506.13131

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 37 AI in Industry

Free notes

Highlights on this page