World Models
Model-free RL buys competence with real interactions, and a robot cannot afford them. A world model spends that same budget learning a simulator instead, then lets the agent practise inside its own imagination.
A world model is a learned simulator. Give it a state and an action and it predicts the next state and the reward. With one in hand, an agent can practise without touching the real environment.
The mechanism that matters is imagination. Compress observations into a small latent state, roll the dynamics forward in that space, and train the policy on trajectories that never happened.
The thing that breaks is horizon. Every imagined step feeds the model its own output, so error compounds. Roll out far enough and the policy is optimising against a fantasy.
01The trade you are making
Look at the two loops first. Every decision further down this page is a consequence of choosing the second one.
Model-free RL, the subject of page 15, learns a policy directly from reward. It is simple and it works. The cost is that every bit of improvement is paid for in real interactions.
That bill is trivial in some places and ruinous in others:
- An Atari emulator runs faster than real time and costs nothing but electricity, so you can spend 200 million frames.
- A physics simulator is nearly as cheap, but it only simulates what somebody wrote down.
- A robot arm runs at one times real time, wears out, and breaks things. Ten million steps would take about a decade, which is not a workable budget.
- A power grid or a live recommender cannot be explored at all. Bad actions have consequences nobody can take back.
Model-based RL changes what the real interactions buy. Rather than spending them on trial and error, spend them on fitting a model of the dynamics. Then do the trial and error inside the model, where a step costs a matrix multiply.
Nothing here is free, because you have swapped a sampling problem for a modelling problem. Model-free RL needs a lot of data but stays honest, since it is always evaluated against the real environment.
A world model needs far less data and is systematically wrong. Its errors are not noise that averages out. They are a bias, and the policy will find it and exploit it. Containing that is what the rest of this page is about.
02What a world model is
Strip the neural networks away and three functions are left. Everything from Dyna in 1991 to Genie is a choice about how to represent them.
- Transition. Given a state and an action, predict the next state.
- Reward. Given a state, predict the reward. Usually learned, occasionally known in closed form.
- Continuation. Predict whether the episode ends here. Easy to forget, and imagined rollouts run straight off the end of the world without it.
A world model is one learned function standing in for the environment's step method.
Dyna, from Sutton in 1991, is the minimal version. Act in the real environment, store each transition, and update the value function on both real transitions and transitions replayed from the learned model. The model manufactures extra experience at no interaction cost.
The payoff is sample efficiency, and it is easiest to see on something small enough to reason about completely.
Try it Two learners, identical data, and the step count each one needs
import numpy as np
rng = np.random.default_rng(0)
N = 20 # a 20-state chain, reward at the far end
def step(s, a): # a=1 moves right, a=0 moves left
s2 = min(N - 1, s + 1) if a else max(0, s - 1)
return s2, float(s2 == N - 1)
Q = np.zeros((N, 2)) # model-free: one backup per real step
T, R = np.full((N, 2), -1), np.zeros((N, 2)) # model-based: keep s', r
s = mf = mb = 0
for t in range(1, 100000): # the SAME random data feeds both
a = int(rng.integers(2))
s2, r = step(s, a)
Q[s, a] += .5 * (r + .95 * Q[s2].max() - Q[s, a])
T[s, a], R[s, a] = s2, r
if not mf and (Q[:-1].argmax(1) == 1).all(): mf = t
if not mb and (T[:-1] >= 0).all():
V = np.zeros(N)
for _ in range(100): V = (R + .95 * V[T]).max(1) # costs 0 steps
if ((R + .95 * V[T]).argmax(1)[:-1] == 1).all(): mb = t
s = 0 if r else s2
if mf and mb: break
print("Q", Q.shape, " learned model T", T.shape, "R", R.shape)
print("model-free policy correct after", mf, "real steps")
print("model-based policy correct after", mb, "real steps")
1337 real steps before its greedy policy points right everywhere, because one real step buys exactly one backup and the reward has to seep back down twenty states. The model-based learner is finished at 145: the moment T and R cover every state-action pair, value iteration extracts the optimal policy with no further interaction at all. The model did not know more. It re-used what it had.The model-based learner needed no extra exploration, no better algorithm and no bigger network. It only had to stop discarding each transition after using it once.
03Compounding error, and where the bill arrives
A model that is 99% right per step is nowhere near 99% right over a rollout, and this failure mode shapes most of the field.
Rolling a model out means feeding it its own predictions. At step two the input is already slightly wrong, so the output is wronger. By step three the input is wronger still. The trajectory walks off the distribution the model was fitted on.
It is the same mechanism that makes behavioural cloning drift on page 21, with the dynamics model in the policy's place.
Try it Fit a near-perfect one-step model, then watch a long rollout leave reality
import numpy as np
rng = np.random.default_rng(0)
# Ground truth: a gentle rotation. The model never sees this matrix.
A = np.array([[0.995, 0.100], [-0.100, 0.995]])
X = rng.normal(size=(50, 2)) # 50 states, short rollouts
Y = X @ A.T + 0.02 * rng.normal(size=(50, 2)) # one noisy step each
Ahat = np.linalg.lstsq(X, Y, rcond=None)[0].T # the entire world model
print("pairs", X.shape, "-> model", Ahat.shape)
print("worst one-step entry error %.4f" % np.abs(Ahat - A).max())
true = model = np.array([1.0, 0.0])
for h in range(1, 201):
true = A @ true
model = Ahat @ model # imagine on, never corrected by data
if h in (1, 5, 15, 50, 100, 200):
drift = np.linalg.norm(model - true)
print("horizon %3d drift %.3f %3.0f%% of true state"
% (h, drift, 100 * drift / np.linalg.norm(true)))
0.0065, which any one-step validation loss would happily call excellent. Fifteen steps in, drift is still only 7% of the state, which is precisely why short imagined rollouts are usable. By horizon 200 the imagined state is 0.972 away from a true state whose norm is about 1, so the prediction carries no information. Nothing is broken: the model is being asked to extrapolate 200 steps from a fit that was only ever checked on one.For deterministic dynamics with Lipschitz constant $L$ and a model whose one-step error never exceeds $\varepsilon$, the gap after $H$ steps obeys a standard bound:
When $L \approx 1$ this collapses to $\varepsilon H$, linear and survivable, which is the benign case above where drift grew almost exactly linearly, at about 0.0048 per step. Any expansion in the dynamics pushes $L$ above 1, the sum turns geometric, and the usable horizon collapses.
The obvious move is to imagine far ahead, since long trajectories say more about the future, but it backfires: value estimates taken from a drifted rollout are confidently wrong, and the policy learns to walk toward wherever the model hallucinates reward.
MBPO made the fix explicit: start every imagined rollout from a real state drawn out of the replay buffer, and keep it short. Many short branches anchored to reality beat a few long fantasies.
The whole Dreamer line follows this rule: imagine briefly and constantly, and re-anchor on fresh real data all the time.
That gives the trade: real steps buy a model, the model manufactures cheap experience, and that experience is trustworthy only over a short horizon.
The rest of the page is how this idea survives contact with pixels. In every successful case the answer is the same: stop predicting pixels.
04Latent dynamics: predict in a space built for predicting
If the observation is a 64×64 colour image, predicting the next one means predicting 12,288 numbers, almost none of which matter for control.
The latent dynamics idea splits the problem in two. An encoder maps an observation to a compact state. The dynamics are learned inside that compact space. A decoder exists mainly to supply a training signal, and planning never calls it.
Three things improve at once. The model has far fewer parameters, so it fits from less data. A rollout becomes a small matrix multiply and stops being a generated image. And the latent is trained to be predictable, which raw pixels never were.
Try it The same system modelled in 64 dimensions and in 2, rolled out 20 steps
import numpy as np
rng = np.random.default_rng(0)
D, K, N = 64, 2, 100 # 64-dim observations of a 2-dim world
A = np.array([[0.99, 0.12], [-0.12, 0.99]]) # true latent dynamics
W = rng.normal(size=(K, D)) / np.sqrt(K) # latent -> observation
Z = rng.normal(size=(N, K))
O = Z @ W + 0.05 * rng.normal(size=(N, D))
On = (Z @ A.T) @ W + 0.05 * rng.normal(size=(N, D))
# (a) predict the observation itself: D*D weights from N pairs
Mpix = np.linalg.lstsq(O, On, rcond=None)[0].T
# (b) find a K-dim latent with an SVD, then predict inside it
E = np.linalg.svd(O, full_matrices=False)[2][:K]
Mlat = np.linalg.lstsq(O @ E.T, On @ E.T, rcond=None)[0].T
print("pixel model", Mpix.shape, " latent model", Mlat.shape)
z = rng.normal(size=K) # a fresh start, unseen in training
pix, lat = z @ W, (z @ W) @ E.T
for h in range(1, 21):
z, pix, lat = A @ z, Mpix @ pix, Mlat @ lat
if h in (1, 5, 20):
t = z @ W
print("h=%2d pixel err %.3f latent err %.3f"
% (h, np.linalg.norm(pix - t) / np.linalg.norm(t),
np.linalg.norm(lat @ E - t) / np.linalg.norm(t)))
h=1 the two models look comparable, 0.012 against 0.008, so a one-step metric would not separate them. The pixel-space model holds (64, 64) free parameters fitted from 100 pairs, so it absorbs the observation noise, and by h=20 its error is 0.490 against the latent model's 0.016. The latent model is (2, 2), four numbers in all, so its advantage is fewer things to get wrong, compounded fewer times.Real systems learn the encoder instead of taking an SVD, and the latent carries a hundred or so dimensions instead of two, but the argument is the same: a learned bottleneck makes long-horizon prediction tractable.
05PlaNet, the RSSM, and the Dreamer line
One architecture family has carried this idea for six years. It helps to know what each version added and why.
PlaNet (2018) learned latent dynamics from pixels and planned directly in the latent space. Its lasting contribution is the recurrent state-space model (RSSM), which keeps two things in the state.
- A deterministic path, a recurrent hidden state that carries information forward reliably. Without it the model forgets what it just saw.
- A stochastic path, a sampled latent that represents genuine uncertainty. Without it the model cannot say "the ball might bounce either way" and averages the two futures into mush.
- h_t is the deterministic state. It uses plain recurrence with no sampling, so information survives many steps.
- q is the posterior. It sees the actual observation $o_t$, so it is only available while training.
- p is the prior. It must guess $z_t$ from $h_t$ alone. Training pulls the prior toward the posterior.
The last line is the trick. Once the prior can predict the posterior without seeing the next frame, you can run the model forward with no observations at all, which is what imagination means here.
Dreamer (2019) changed what you do with the model. PlaNet ran a fresh search at every step. Dreamer instead trains an actor and a critic on imagined latent trajectories, backpropagating value gradients through the learned dynamics. An expensive online search is replaced by a network that has already done it.
- Dreamer (2019) learns an actor-critic purely inside latent imagination, on visual control tasks.
- DreamerV2 (2020) uses discrete categorical latents in place of Gaussian ones, which handle multimodal futures better. It was the first agent to reach human-level on the 55-game Atari benchmark by learning behaviours inside a separately trained world model.
- DreamerV3 (2023) adds normalisation and transformation fixes that make a single hyperparameter setting work across more than 150 tasks. Applied out of the box, it was the first system to collect diamonds in Minecraft from scratch, with no human data and no curriculum.
- DayDreamer (2022) runs the same algorithm online on physical robots. A quadruped learned to roll off its back, stand and walk in one hour of real experience, without a simulator and without resets.
DayDreamer is the cleanest evidence for the claim at the top of this page. One hour of wall-clock experience is the kind of budget a physical robot has.
06Planning inside the model
A world model hands you a cheap simulator. What you do with it is a separate choice, and there are three standard answers.
- Sample and score. Draw hundreds of action sequences, roll each through the model, keep the best few, refit the sampling distribution, repeat. This is the cross-entropy method and needs no policy network. PlaNet and PETS work this way.
- Train a policy in imagination. This is Dreamer's answer: amortise the search into an actor and a critic so that acting is one forward pass.
- Search with a learned value. Monte Carlo tree search over a learned model, which is MuZero.
MuZero is the interesting case, because it declines to model the world. Its learned dynamics predict three quantities and no others: reward, value and policy. It never reconstructs a frame or a board position.
This is value equivalence: the model has to be right about the quantities the search consumes, and any detail that cannot change a value is free to be wrong.
Every practical planner replans: it computes a plan of length $H$, executes one or two steps, observes what happened, then plans again from the true state. This is model predictive control.
Replanning resets the compounding error to zero on every step. A plan that would be nonsense at horizon 50 is perfectly fine when you only ever execute its first action.
The same machinery gives you counterfactual prediction. The model answers "what would have happened had I turned left" without turning left, which is how an agent compares options it never took. Generating counterfactuals is cheap, but being right about them is hard.
So the control half of the field combines latent dynamics, imagination, and planners that replan often enough to survive their own model error.
The other half arrives from a completely different direction. It starts with video generation and works backwards toward control.
07Video prediction as a world model
A model that predicts the next frame of a video is a world model with no actions and no reward. Adding the actions turns it into something a policy can use.
Action-conditioned video prediction conditions the next frame on the past frames and the action taken. Finn and Levine did this for robot arms in 2016 under the name visual foresight, then planned pushes by searching over action sequences and scoring the predicted frames.
The modern versions are generative video models, and they are far better at frames:
- GameNGen (2024) trained a diffusion model to predict the next DOOM frame from past frames and inputs. It runs at 20 frames per second on a single TPU, and human raters are barely better than chance at telling short clips from the real game.
- Genie (2024) learned a latent action space from unlabelled internet video. It can turn a single image or sketch into a controllable environment, with no action labels anywhere in training.
- GAIA-1 (2023) and NVIDIA's Cosmos platform do the same for driving and physical AI, generating futures conditioned on actions and text.
A simulator is written by a person. It is consistent by construction and wrong wherever that person's physics was wrong. You can query it anywhere, including states no data ever covered.
A generative video model produces plausible frames, and plausible frames are not always consistent: an object can leave the frame and come back a different colour, and nothing in a per-frame loss objects.
A world model is judged by whether a policy trained inside it works in the real environment. That test is much harder than looking right.
The live question through 2025 and 2026 is whether the video route reaches control. Frame quality and real-time interactivity have improved enormously. Long-horizon consistency and physical reasoning have improved much less, and those are exactly the properties a planner leans on. Page 26 tracks where that has got to.
08JEPA: predict representations, not pixels
There is a third position worth taking seriously: predicting pixels is the wrong objective in the first place.
A pixel loss spends most of its capacity on things that are both unpredictable and irrelevant. Leaf texture, sensor noise, the exact shape of a shadow. A model that gets those wrong is penalised as hard as one that misses the oncoming car.
A joint-embedding predictive architecture (JEPA) predicts in representation space instead. Encode the context, encode the target, and train a predictor to match the target's embedding. There is no decoder and no reconstruction.
The obvious objection is collapse. If the encoder emits a constant, prediction is perfect and nothing has been learned. JEPA variants avoid this with asymmetry: a stop-gradient, a slowly-updated target encoder, or both. Page 03 covers those tricks in their original setting.
V-JEPA 2 (2025) carried this into control. Pre-train on over a million hours of internet video with no action labels, then post-train an action-conditioned latent model on under 62 hours of unlabelled robot video, and plan with it zero-shot on a real Franka arm.
The bet is to learn the structure of the world from video nobody labelled, then attach actions using a small amount of interaction data.
09Evaluating a world model
The metric you reach for first is the one most likely to mislead you.
A world model has two very different jobs, and they are measured in incompatible ways:
- Prediction quality. One-step and multi-step error, PSNR, FVD for video. Cheap, reproducible, and only loosely related to usefulness.
- Downstream control utility. Does a policy trained or planned inside the model perform in the real environment? Expensive, noisy, and the only thing that settles the question.
Reconstruction loss is dominated by whatever occupies the most pixels. In a driving scene that is road and sky. In a manipulation task it is the tabletop and the wall behind it.
A model can lower its loss by rendering the background more precisely while getting the small, fast, control-relevant object wrong. The numbers improve and the policies get worse. This is why PlaNet argued for judging the model on predicted reward over several steps and not predicted observations.
Three specific things to probe, because an aggregate error number hides all of them:
- Long-horizon consistency. Roll out for hundreds of steps and check the scene is still the same scene. Does the room still have the same number of doors?
- Object permanence and physical reasoning. Occlude something, wait, reveal it. Benchmarks such as IntPhys and Physion probe this directly, and models that look excellent frame to frame fail it routinely.
- Counterfactual accuracy. Take one start state and two different action sequences. The model has to predict different futures, and not a blurred average of both.
One more catches teams late: error where the policy goes. A model fitted on data from the old policy is accurate on the old policy's states. Training a new policy inside it moves the state distribution to exactly the region the model was never fitted on.
10Build this
The horizon trade-off on this page is a curve with a peak in it. You can plot that curve in an afternoon, on a laptop, in numpy.
Build a control task small enough to solve exactly, learn its dynamics from a handful of real episodes, then train policies inside the model at several imagination horizons. What you are hunting for is a curve that rises and then falls.
- Write a 2-D point-mass environment in numpy. State is position and velocity, action is a bounded acceleration, reward is negative distance to a target. Solve it with a scripted controller first so you know what good looks like.
- Collect 20 real episodes under a random policy. Fit a linear dynamics model on the one-step transitions. Report its one-step error and its 50-step rollout error as two separate numbers.
- Train a policy by random shooting inside the learned model, at imagination horizons H of 1, 3, 5, 15 and 50. Evaluate each resulting policy in the real environment, never in the model.
- Plot real-environment return against H. Record where the peak sits, and keep the imagined return on the same axes for comparison.
- Now alternate: collect 5 more real episodes using the current policy, refit the model, retrain, repeat five times. Watch the peak move to the right as the model improves.
- The breakage: freeze the model after step 2, never refit it, and push H to 200.
11What breaks
World-model failures are rarely a bad architecture. They are a model being trusted somewhere it was never fitted.
- The policy finds a hole in the model. Imagined return climbs steadily while real return is flat or falling. The agent has located a state where the model promises reward the environment does not pay.
- One-step error looks superb and 50-step rollouts are garbage. Validating on one-step transitions is the default, and it certifies nothing about rollouts.
- The model is accurate everywhere the old policy went. Then the new policy goes somewhere else, and the model has no data there. Each round of improvement invalidates part of the model.
- Imagined episodes never end. The continuation head was skipped, so the agent imagines collecting reward past a terminal state and learns to run into walls.
- The stochastic latent collapses. The posterior stops carrying information, the model turns deterministic, and every uncertain future is predicted as its own blurry average.
- Long rollouts lose the scene. Objects change colour, doors vanish, counts drift. A video world model can be excellent frame to frame and no longer the same world after thirty seconds.
- Reward is the hardest head to fit and the one that matters most. With sparse reward almost every imagined transition has reward zero, so the head predicts zero everywhere and imagination becomes worthless.
- Weeks go into the model's own metrics. Reconstruction loss falls steadily and control performance does not move, because nobody was running the policy in the real environment.
12Where you meet this in the wild
Same idea, four very different reasons for wanting it.
This is the original motivation. Real steps are slow, wear out joints and break objects, so the only viable budget is hours of interaction and not years.
Latent world model trained online, short imagined rollouts, constant refitting.
Used mostly to manufacture the rare events that logged data does not contain. Generate the cut-in, the pedestrian at dusk, the debris in lane, then test a planner against them.
Scenario generation, counterfactual replay and closed-loop evaluation.
Not, so far, the thing steering the car.
Here the world model is the product. A learned engine can be conditioned on an image or a sentence, which no hand-built engine can do, and it generates training environments for agents at scale.
Action-conditioned video generation with a latent action space.
Data-centre cooling, chemical plants, grid dispatch. Exploration is not permitted, so a model is fitted from logs and a controller plans against it. Much of this is system identification wearing newer clothes.
Learned dynamics plus MPC, with hard safety limits enforced outside the model.
13Interview questions
BeginnerWhat is a world model, and why would you train one?
A world model is a learned simulator of an environment's dynamics: given a state and an action it predicts the next state, the reward, and usually whether the episode has ended. You train one because model-free reinforcement learning pays for every improvement in real interactions, and in many settings those interactions are the scarce resource. A game emulator makes them nearly free, but a robot arm runs at one times real time and wears out, and a power grid or a live production system cannot be explored at all. A world model redirects the interaction budget: instead of spending real steps on trial and error, you spend them on fitting the dynamics, and then do the trial and error inside the model where a step costs a matrix multiply. The gain is sample efficiency and the cost is that the model is systematically wrong in ways the policy will find.
BeginnerWhat is the difference between model-based and model-free RL?
Model-free methods learn a policy or a value function directly from experienced transitions and never represent the environment's dynamics explicitly. Q-learning, DQN and PPO are all model-free. Model-based methods learn an explicit transition function, and often a reward function, then use it either to generate extra training data or to plan by searching over imagined action sequences. The practical difference is where the samples come from. A model-free learner performs one update per real transition, so improvement is bounded by how fast the environment can be run. A model-based learner can replay the same real transitions through a learned model to manufacture many more, or plan without collecting anything new. Model-free is asymptotically more reliable because it is always judged against reality; model-based is dramatically more sample efficient and inherits every error in the model it learned.
IntermediateWhy does compounding error limit rollout length, and what do you do about it?
Rolling a model forward means feeding it its own output, so the input at step two already contains the step-one error, and each subsequent prediction is made from a state further off the training distribution. Errors therefore accumulate rather than cancel. For deterministic dynamics with Lipschitz constant L and one-step error at most epsilon, the horizon-H gap is bounded by epsilon times (L^H minus 1) over (L minus 1), which is linear in H when L is near one and geometric as soon as L exceeds one. The consequence is that a model with an excellent one-step validation loss can produce a rollout that is completely uninformative fifty steps out. The standard mitigations are to keep imagined rollouts short, to branch each of them from a real state sampled out of the replay buffer rather than from an imagined one, which is MBPO's contribution, to replan from the true state after executing only one or two actions, which is model predictive control, and to keep refitting the model on fresh data as the policy changes where it goes.
IntermediateWhat does the RSSM's split into deterministic and stochastic state buy you?
Each path fixes a failure of the other. A purely stochastic latent is resampled at every step, so information has to survive repeated sampling to persist, and in practice the model forgets context quickly and long-horizon prediction degrades. A purely deterministic recurrent state remembers reliably but cannot represent genuine uncertainty, so when the future is multimodal it predicts the average of the modes, which is often a state that cannot physically occur. The RSSM carries both: a recurrent deterministic hidden state that transports information across many steps, and a sampled stochastic latent conditioned on it that expresses what is genuinely unknown. Training fits a posterior that sees the current observation and a prior that does not, and pulls the prior toward the posterior. That is what makes imagination possible, because at rollout time only the prior is available and it has been trained to stand in for the posterior without ever seeing a frame.
IntermediateWhy does Dreamer learn behaviours in latent space rather than in pixel space?
Three reasons, all of them practical. First, cost: an imagined step in a compact latent is a small matrix operation, whereas an imagined step in pixel space means generating an image, and Dreamer needs enormous numbers of imagined steps to train an actor and a critic. Second, fit: a latent dynamics model has orders of magnitude fewer parameters than a pixel-space predictor, so it can be fitted from the small amount of real data that made the whole approach worthwhile. Third, and most importantly, predictability: the encoder is trained so that the latent is something the dynamics can actually predict, while raw pixels contain a great deal of high-frequency detail that is both unpredictable and irrelevant to control. A pixel loss forces capacity onto texture and noise. The decoder exists to supply a training signal for the representation, and planning never needs to call it.
DeepMuZero never predicts observations. What does it predict, and why is that enough?
MuZero learns three functions: a representation function that maps the observation history to an abstract state, a dynamics function that takes an abstract state and an action to the next abstract state plus a predicted reward, and a prediction function that outputs a policy and a value from an abstract state. Nothing reconstructs an observation, and the abstract state has no imposed semantics at all. This is enough because Monte Carlo tree search only ever consumes rewards, values and policy priors, so a model that is correct about those is indistinguishable, from the search's point of view, from a perfect simulator. That is the value-equivalence principle: the model must be accurate about the quantities that affect decisions, not about the world. The benefit is that all of the model's capacity goes into decision-relevant structure rather than into rendering, which is why it works on Atari pixels and on Go without being given the rules. The cost is that the model is not interpretable or reusable for anything other than the task whose values it was trained on.
DeepYour world model's validation loss is excellent but policies trained in it fail. Diagnose.
The first suspicion is that validation is measuring the wrong thing. One-step prediction error on held-out logged transitions says nothing about multi-step rollouts, so re-measure as multi-step rollout error and, more importantly, as predicted cumulative reward over the horizon you actually imagine at. The second suspicion is distribution shift caused by the policy itself: the model was fitted on data from an earlier policy, and training a new policy inside it drives the state distribution into regions with no data, where the model is extrapolating and the validation set cannot see it. Look for imagined return climbing while real return stays flat, which is the signature of model exploitation, and check whether the visited latent states during imagination fall outside the training set's support. Other common causes are a missing or badly fitted continuation head, so imagined episodes never terminate, a collapsed stochastic latent that averages multimodal futures, and a reward head that predicts zero everywhere because reward is sparse. The fixes are structural rather than clever: shorten the imagination horizon, branch rollouts from real replayed states, use an ensemble and penalise disagreement so the policy is discouraged from going where the model is uncertain, and refit the model on freshly collected on-policy data every iteration.
DeepWhen would you not use a world model?
When real interactions are cheap and the dynamics are hard to model, which is exactly the regime where model-free methods win. If you have a fast deterministic simulator, running it is both free and exactly correct, so learning an approximation of it buys nothing and introduces bias; you are better off spending the compute on more environment steps. The same argument applies when a good hand-written simulator already exists and the deployment gap is one of appearance rather than physics, where domain randomisation is usually the cheaper tool. World models also struggle where dynamics are chaotic or dominated by other agents, since prediction error grows too fast for any usable horizon, and in highly stochastic environments where the model must represent a wide distribution rather than a trajectory. A final practical case against is engineering cost: a world model adds an encoder, a dynamics model, reward and continuation heads, and an imagination loop, all of which can fail quietly. If you can afford the samples, the simpler system is usually the right call.
14Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.