←Home KnowML
Systems, Safety & InterviewChapter 26

2026 Frontier Map

Every technique on this site up to here answers "how do we build a good model." This page is about a quieter shift. Most of the field's energy in 2025–2026 went into what you do around a model, at inference time and as a system, and less into making the model itself bigger.

18 min read Assumes: LLM training (10), RAG & agents (11), efficient systems (23)
Start reading
TL;DR

Pretraining-scale gains got expensive faster than they got useful. So the frontier moved to spending compute at inference time instead: sample several candidate answers, score them with a verifier, keep the best one. That can beat a much larger model on the same problem for less total cost.

The same shift shows up as agentic systems. Instead of asking one model to do everything in one forward pass, wrap it in a loop that plans, calls tools, checks its own work, and iterates. You trade a monolithic black box for an inspectable pipeline built from pieces that are each individually reliable.

Multimodality, robotics and science followed the same arc: bolt a pretrained backbone onto a new modality or embodiment instead of starting from scratch. None of this solved the field's hard problems: reliability at the tail, grounding claims in evidence, planning over long horizons, and evaluation itself all remain open.

It just gave everyone new tools to attack them with. One sentence for an interview: the 2026 frontier is less "a bigger model" and more "a system built around a fixed model," and the honest caveat is that the open problems from 2023 are mostly still open.

01Intuition

Pretraining spends the same compute on "what's the capital of France" as on "prove this theorem." That is the problem.

Every large model has a fixed amount of knowledge and skill baked in by pretraining and fine-tuning. Call this raw competence. Historically the way to get more of it was to make the model bigger, train on more data, or both, all before it answers a single question.

That has a ceiling problem. Pretraining compute is a sunk, fixed cost. The model cannot tell in advance which questions deserve more effort, because by the time it is answering, all the training compute has already been spent.

The inversion the whole page rests on

Inference-time compute flips the order: you spend a little more per question, adaptively, right when you know what the question is, and less upfront.

You can sample multiple attempts, search over partial solutions, or run a verifier to catch mistakes.

An easy question gets one quick pass. A hard one gets ten attempts and a checker.

Agentic systems apply that same "spend adaptively" idea one level up. Rather than forcing a single model call to plan, execute and self-correct all at once, split those into separate steps a system can run in a loop. Stop early on easy tasks. Iterate longer on hard ones.

Multimodal and embodied frontier work is a different kind of adaptation: reusing a capable pretrained backbone for a new input or output modality, like video or robot actions, instead of learning that modality from zero. The expensive part, broad world knowledge and reasoning, transfers. Only the comparatively cheap part, how to perceive or act in the new modality, needs learning fresh.

One way to get more from a model at test time is to ask it the same question many times and take the majority answer. Suppose each attempt is right 60% of the time and the attempts are independent.

Try it Take a majority vote over repeated attempts
import numpy as np

rng = np.random.default_rng(0)
p = 0.6                                   # chance one attempt is right

for k in (1, 5, 15, 51):
    votes = rng.random((100_000, k)) < p  # k independent attempts
    majority = votes.sum(axis=1) > k / 2
    print(f"{k:>2} attempts, majority vote: right {majority.mean():.1%}")
One attempt is right 60% of the time. The majority of 5 is right about 68% of the time, of 15 about 79%, and of 51 about 93%. The gain holds only if the mistakes are independent. Attempts from the same model can fail in the same way, so real gains are smaller than this idealised case.

02Timeline

  • Chain-of-thought prompting (2022) showed that asking a model to think step by step improved accuracy on multi-step problems with no architecture change, which raised the question of spending more compute on that scratch space and picking the best result. That became test-time-compute scaling and process-supervised verifiers (2023–2024).
  • In parallel, RLHF and instruction tuning (covered in 10) made models reliable enough at following instructions and calling tools to wire into agentic loops that act, observe a result and re-plan.
  • This page covers five directions the field pushed on, each a response to a specific limitation of what came before, plus the problems none of them fixed. Page 11 covers the retrieval and single-step tool-use side.

03The landscape, in five directions

Reasoning & inference-time compute

The clearest instance of "spend compute per-query instead of upfront." Snell et al. formalize two mechanisms (see Go Deeper):

  • Search against a verifier. Sample N candidate solutions, score each with a model trained to judge correctness, keep the best.
  • Adaptive revision. The model iteratively refines its own draft, conditioned on the prompt.

The verifier is usually a process reward model (PRM), scoring each intermediate reasoning step as well as the final answer. Lightman et al.'s "Let's Verify Step by Step" showed the difference: process supervision caught errors that outcome supervision let slip through, because a right final answer can hide a wrong method that happened to land correctly.

Agentic systems

Page 11 already covers the core loop: reasoning, a tool call, an observation, repeat (ReAct). This page's addition is what changes when that loop runs for a long horizon. Picture a coding agent planning a multi-file change, writing code, running the test suite, reading the failure and retrying, possibly dozens of times.

That demands three things the single-step loop does not:

  • Memory. It carries state across steps.
  • Planning. Break the goal into sub-goals.
  • Self-verification. Decide for itself when a sub-goal succeeded, which is more than running without crashing. This is the verifier idea from reasoning, one level up.

Toolformer (Schick et al., see Go Deeper) was an early demonstration that a model can learn when to call a tool, and with what arguments, from self-supervised examples instead of hand-scripted rules. This underlies how modern agent frameworks decide to reach for a calculator, a search engine, or a code interpreter.

The multimodal & efficiency frontier

Two threads, both pushing a capable backbone into new territory.

  • Multimodal. Page 14 covers the bolt-on "encoder + projector + LLM" recipe versus training natively multimodal from the start. The frontier direction pushes native multimodality across more modalities at once (text, image, audio, video in one model), and toward generative video temporally consistent enough to function as an interactive world simulator rather than a short clip generator. A direct extension of the diffusion techniques in 12, applied across time as well as space.
  • Efficiency. Mixture-of-Experts (Switch Transformers, see Go Deeper, covered in depth in 10) decouples total parameter count from per-token compute, so capacity can grow without inference cost growing at the same rate. Distillation and quantization (23) push the other way: take a large model's behavior and compress it small enough to run on-device.

Both chase the same goal, which is more capability per dollar of inference compute and less reliance on training compute.

The embodied & robotics frontier

Vision-language-action models (21) and world models used as learned simulators for planning (22) are the same "reuse a pretrained backbone" idea, applied to physical action instead of text or code. You start from a model that already understands images and language and fine-tune it to output robot actions as an additional token type, which beats training a control policy from scratch on a small amount of robot data.

Cross-embodiment generalization (21) pushes further, pooling data across many different physical robot bodies so a single policy transfers across them. The same logic that makes a pretrained language model reusable across many downstream text tasks.

AI for science

The clearest proof that a general architecture, plus the right training objective and enough of the right data, generalizes past language and images.

AlphaFold2 (Jumper et al., see Go Deeper) treated protein structure prediction as a geometric learning problem and reached accuracy that decades of physics-based simulation had not. It became Nature's and Science's Method of the Year for 2021 and, per the community assessments cited in follow-up literature, changed how structural biologists work day to day.

The same recipe of domain-appropriate architecture, a large corpus of structured natural data, and a well-chosen objective has since been extended toward molecule and materials design. Those domains are individually less mature than protein structure prediction, and remain active, unsettled research areas.

04The core mechanisms, precisely

This page spans too many sub-fields to have one shared equation the way a single-architecture page does. But two mechanisms recur across almost everything above, so they are worth stating exactly.

Best-of-N with a verifier. Given a prompt, sample N candidate completions from the model, independently and usually at temperature > 0 so they differ. Score each with a verifier V and return argmax_i V(candidate_i).

The verifier can be a separate trained model, such as a PRM scoring each reasoning step. In simpler setups it is a rule-based checker: does the code compile, does the final answer match a known constraint.

The core tradeoff is that larger N costs inference compute linearly, while accuracy gains are sublinear and eventually flatten. This is the compute-optimal allocation problem Snell et al. study: for a fixed budget, is it better to spend on more samples, or on letting one sample think longer?

Sparse routing (Mixture-of-Experts). A router network computes, per token, a score over a bank of expert feed-forward networks. Only the top-k highest-scoring experts, commonly k=1 or k=2, process that token. The rest contribute nothing to that token's forward pass and cost nothing for it.

Total parameter count can scale with the number of experts, while compute per token stays pinned to k experts' worth of work. This is why sparse models can hold far more total parameters than a dense model trained at the same per-token compute budget. See 10 for the training-time complications this introduces, like load balancing across experts.

05Why this direction

Obvious alternative

The first path is to keep scaling pretraining with more parameters and more data. That recipe reliably worked from GPT-2 through GPT-4-class models.

→
What frontier work does instead

Spend more compute adaptively, at inference time, per query. A fixed pretraining budget is spent identically on easy and hard questions. Inference compute can be allocated only where it is needed.

→
What it costs

Higher and less predictable per-query latency and serving cost. A hard query might take 10× longer, or run 10× more samples, than an easy one. That complicates capacity planning versus a model with fixed inference cost.

Snell et al.'s headline result makes the tradeoff concrete. Under a matched compute budget, their compute-optimal test-time strategy let a smaller model outperform a roughly 14× larger model, on problems where the smaller model already had some non-trivial baseline success rate.

The caveat to say out loud in an interview

That result depended on the problem having a verifiable or scoreable structure, like math or code, where a verifier or search strategy can tell a better answer from a worse one.

For open-ended tasks with no clean way to score a candidate, spending more inference-time compute has much weaker guarantees.

Which is exactly why verifier design and evaluation (25) are as much a frontier problem as raw compute allocation.

Why agentic decomposition? A monolithic model has to implicitly learn arithmetic, retrieval and code execution as internal skills, imperfectly, from training data. Routing each to an actual calculator, search index or code interpreter does that sub-task exactly. And a wrong tool call is visible and debuggable in a way a wrong internal computation buried inside a forward pass is not.

06Complexity, failure modes, and the persistent open problems

Inference-time search inherits a classic optimization failure mode: reward hacking against the verifier. Every verifier is imperfect to some degree. Aggressive search can find a candidate that scores well under the verifier without being correct, the same way an RL policy can learn to exploit a misspecified reward function (15).

The perverse part: the more compute you spend searching, the more thoroughly you search for exactly this kind of exploit. Whether you meant to or not.

Agentic loops inherit a different one: compounding error over long horizons. Say an agent's per-step success rate is 95% (illustrative number, not a benchmark result) and a task takes 50 sequential steps. If errors are independent, the probability every step succeeds is roughly 0.95^50 ≈ 7.7%.

So long-horizon reliability is the bottleneck for autonomous coding, more than per-step accuracy. Self-verification at intermediate steps matters far more here than in a single-shot model call, because it catches an error at step 12 before it surfaces at step 50.

And then there is the honest list. These problems carried over unchanged from 2023, and none of the directions above solved them:

  • Reliability. Models are still confidently wrong at a non-trivial rate. Inference-time compute reduces this but does not eliminate it.
  • Grounding. Tying a claim to verifiable evidence and not to a fluent-sounding guess. See hallucination and groundedness in 25.
  • Long-horizon planning. The compounding-error problem above, unresolved for long, ambiguous real-world tasks.
  • Evaluation. Benchmark contamination and the weaknesses of LLM-as-judge (25) get worse as models get better at sounding right.
  • Cost. Inference-time compute trades money and latency for accuracy directly. There is no free lunch in that trade.
  • Safety and controllability. An agent that can execute code and call tools has a much larger blast radius when it does the wrong thing than a model that only outputs text.

None of this is a reason for pessimism about the field; it is the current research agenda. Naming these problems specifically gives a more informed answer than gesturing vaguely at "AI safety."

07Build this

Every claim on this page rests on one curve: accuracy against sampling budget. You can measure it yourself in an afternoon, and the shape it comes out is more informative than any summary of it.

Project Measure the gap between finding an answer and picking it ~3 hours · any chat API or local model

Take a set of problems with automatically checkable answers, grade-school math being the standard choice. Sample N answers per problem at non-zero temperature. Then plot two curves on the same axes as N grows: how often any of the N samples is correct, and how often your selection rule picks a correct one.

  1. Pick 100 problems where the answer is a number, so grading is exact string comparison and never a judgement call.
  2. Sample N=16 answers per problem at temperature around 0.8. Cache every sample to disk. This is the only slow step, and you pay for it once.
  3. From that cache, compute both curves for N = 1, 2, 4, 8, 16 by subsampling, with no new API calls.
  4. Curve one is the oracle: was any sample right. Curve two is majority vote: take the most common final answer.
  5. Now change the selection rule. Score each candidate by its own average token log-probability and pick the highest, then plot that as a third curve.
You'll know it worked when the two curves separate and keep separating as N grows. The oracle curve rises; majority vote rises more slowly and starts to flatten. That widening gap is the entire test-time-compute argument in one picture: the model already produced a correct answer, and the thing standing between you and it is selection.
What the third curve teaches. Self-reported confidence is a weak selector, and on some problem sets it does worse than plain majority vote. A model can be fluent and wrong with high probability on every token. That is the argument for training a separate verifier on reasoning steps rather than reading confidence off the generator, and it is why process reward models exist at all.

Where this runs in production

An autonomous coding agent, given a bug report, runs a loop:

  • Plan. Identify likely files, form a hypothesis.
  • Act. It writes a code change.
  • Verify. Run the existing test suite.
  • Revise. If tests fail, read the failure output and try again. Loop until tests pass or a step or time budget is exhausted.

This is the agentic loop from section 03, with the test suite standing in as an automatic, domain-specific verifier. It has the same "generate, verify, keep or retry" shape as best-of-N sampling in section 04, spread across sequential steps instead of parallel samples. And it inherits the compounding-error risk from section 06 whenever a multi-file change needs many sequential correct steps.

08Interview questions

ConceptWhy might a smaller model with more inference-time compute beat a larger model with less?

Because pretraining compute is spent identically on every future query regardless of difficulty, while inference-time compute (sampling, search, verification) can be allocated adaptively, per query, only where it's needed — an easy query doesn't need the extra spend, a hard one benefits from it directly. Snell et al. show this concretely: under a matched total compute budget, adaptive test-time strategies let a smaller model outperform a much larger one on problems with verifiable structure. The caveat worth stating: this depends on having a way to actually score candidate answers (math, code) — it's a much weaker effect on open-ended tasks without a clean verifier.

System designWhy does an agent's reliability degrade so much faster than its per-step accuracy would suggest?

Because end-to-end success requires every step in a long sequential chain to succeed, and independent per-step error compounds multiplicatively — a 95% (illustrative) per-step success rate over 50 sequential steps gives roughly 0.95^50 ≈ 7.7% end-to-end success. The fix isn't a better single model call, it's catching errors at intermediate steps (self-verification) so a mistake at step 12 gets corrected before it propagates to step 50, rather than being discovered only at the end.

Failure modesWhat's "reward hacking against a verifier," and why does more search make it worse, not better?

A verifier is itself an imperfect, learned approximation of "is this actually correct." Aggressive search or sampling explores a huge space of candidates specifically looking for whatever scores highest under that verifier — the more thoroughly you search, the more likely you are to find an edge case where the verifier's score and true correctness diverge, which is a structurally identical failure to reward hacking in RL (15): optimizing hard against an imperfect proxy objective finds exactly the proxy's blind spots.

Honesty checkName two problems from 2023 that are still unsolved in 2026, and say why.

Reliability and evaluation are two honest answers. Reliability: inference-time compute and verifiers reduce error rates but don't eliminate confidently-wrong outputs, because the verifier itself is an imperfect model, not ground truth. Evaluation: LLM-as-judge and benchmark-based evaluation (25) get harder, not easier, as models improve, because a more fluent model is also more capable of sounding correct while being wrong, and benchmarks leak into training data over time (contamination). Naming the actual mechanism behind why a problem is still open reads as far more credible than just listing "safety" or "hallucination" as buzzwords.

09Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 26 2026 Frontier Map

Free notes

Highlights on this page