II · THE IDEA · ARTIFICIAL INTELLIGENCE
Model-Based RL with Learned Dynamics
Written before source checking was enforced · facts in this lesson are unverified
▶ Listen · narrated
Real environment steps are often slow, costly or fragile. If a network can predict the next state well enough, most of the learning can move into simulated rollouts the agent invents for itself.
At a glance
- What it is
- Learn a transition model, then plan or train on imagined trajectories
- Main parts
- Dynamics model, policy or planner, real data buffer
- Hoped gain
- Fewer real environment steps when predictions stay reliable
- Main risk
- Model error compounds over long imagined horizons
Think of learning to ride a bicycle on a quiet street that you can visit only a few minutes a day. You might also sketch the street on paper and move a toy bike through turns you have not tried yet. The sketch is your dynamics model: it guesses what happens if you lean or brake. Practising on the sketch is an imagined rollout. If the sketch matches the real street, those practice runs teach useful balance. If the sketch is wrong about gravel or slopes, you may “learn” tricks that fail the moment you go outside again.
In model-based RL the agent collects real experience, fits a network that predicts the next state (and often the reward) from the current state and action, then uses that network as a stand-in world. It can plan a few steps ahead before acting, or train its policy on thousands of made-up trajectories between rare real episodes. The real world stays in the loop to correct the sketch; the sketch exists to multiply practice when real steps are scarce.
Model-based RL with learned dynamics maintains at least three pieces: a dataset D of real transitions (s, a, r, s′), a parametric dynamics model pθ(s′, r | s, a) (or a latent transition model when observations are partial), and a policy or planner π.
Training alternates. (1) Act in the true MDP with π, perhaps with exploration noise, and append transitions to D. (2) Fit θ by supervised or likelihood training on D; ensembles or probabilistic heads are often used so that disagreement or variance can stand in for uncertainty. (3) For K imagination steps, sample start states from D or from a learned prior, then roll out actions from π (or from a shooting/MPC planner) through pθ, producing synthetic trajectories τ̂. (4) Update π with a model-free-style objective on τ̂—policy gradients, Q-learning, or return-weighted regression—or, at decision time only, choose a = first action of a short-horizon plan that maximises predicted return under pθ, then replan after the real s′ arrives.
Limitations are sharp. Autoregressive rollouts compound local error; the effective horizon is often far shorter than the task horizon. Model bias becomes policy bias when π optimises against systematic errors in pθ (reward hacking in imagination). Covariate shift appears as soon as π leaves the state-action support of D. Mitigations include short horizons, aggregation of real on-policy data, uncertainty-aware pessimism, latent models that predict only what the controller needs, and always scoring candidates by real return rather than imagined return alone.
Implementations differ mainly in where pθ is allowed to influence π: background batch training on imagined data, online MPC with a terminal value function, or both. The shared structure is outer real data collection and inner use of a learned simulator for planning or policy improvement.
Look closer
Two nested loops, not one
Real interaction still happens, but it is no longer the only source of learning signal. The outer loop collects genuine transitions and refits the dynamics model. The inner loop rolls the current model forward from real or sampled states, producing synthetic trajectories on which a policy can be improved or a planner can search. The quality of the whole scheme hangs on how faithfully those inner rollouts track the true environment—and for how many steps they do so before drift dominates.
Where the model is allowed to speak
Some designs use the learned dynamics only for short-horizon planning at decision time, discarding the imagined futures once an action is chosen. Others treat imagination as a training substrate: thousands of synthetic episodes update the policy while the real environment is touched sparingly. A third pattern mixes both. Each choice changes how model bias leaks into behaviour. Planning that re-anchors on the true state every step is often more forgiving of an imperfect model than long unsupervised rollouts used as if they were data.
Compounding error is the visible failure mode
One-step prediction can look strong in isolation and still fail as a simulator. Small local mistakes in next-state prediction accumulate; after tens of steps the imagined trajectory may occupy regions the model never saw in real data, where its outputs become confident and wrong. Policies trained there can exploit those fantasies. Practitioners therefore watch horizon length, uncertainty estimates, and the gap between returns in imagination and returns on real rollouts—not only average one-step loss.
The story
Model-based reinforcement learning with learned dynamics starts from a simple shortage. Interacting with the true environment may be slow, expensive, unsafe, or limited by hardware. Model-free methods learn a policy or value function from those interactions alone. Model-based methods try to spend some of the budget on building a simulator—usually a neural network that maps a state and action to a predicted next state and, often, a reward—and then to extract more learning from that simulator than from the raw stream of real experience.
The dynamics model is fit on a replay of real transitions. In the fully observed case the training target is ordinary supervised prediction: given (s, a), forecast s′ and r. When the observation is partial, the model may be a latent-state sequence model that maintains a belief and predicts forward in that latent space. Either way, once the model exists, the agent can sample imaginary trajectories by feeding its own predicted states back in as inputs, choosing actions from the current policy or from a planner.
Those imagined trajectories can be used in several distinct ways. A planner may search over action sequences under the model and execute only the first action before replanning from the next real observation. Separately, or instead, a policy gradient or value-based update can treat the synthetic rollouts as if they were experience, updating the controller many times between real episodes. Hybrid schemes do both: imagination for bulk credit assignment, and short model-based look-ahead at decision time.
None of this removes the need for real data. The model is only as good as the transitions it has seen, and policies that leave the support of that data expose the model’s blind spots. A practical training loop therefore alternates: act in the world (often with exploration), store transitions, improve the dynamics fit, then run a burst of imagined rollouts to improve the policy or the plan. The real environment remains the final judge; imagined return is a proxy that can diverge.
The central engineering tension is horizon. Longer imagined rollouts give the policy richer multi-step signal and let a planner reason further ahead, but each extra step multiplies model error. Short rollouts stay closer to truth yet may not reach the consequences that matter. Uncertainty-aware models, ensembles, and truncation of imagination when disagreement grows are common responses to that tension. So is anchoring: whenever a real observation arrives, discard the imagined state and start again from the truth.
What one actually implements, then, is not a pure substitute for the world but a carefully limited second source of experience. The learned dynamics are a neural simulator used for planning and for training with imagined rollouts—powerful when local prediction is solid and the agent keeps checking back with reality, and brittle when imagination is trusted past the point the model can support.
Why it mattered then
As deep reinforcement learning moved into domains where each real step carried a measurable cost—robotics hardware, long game episodes, constrained simulators—the appeal of squeezing more learning from each transition grew. Learning a dynamics model offered a route to sample efficiency that did not require a hand-built physics engine. The same neural machinery used for policies could, in principle, be turned into a general predictor of what happens next, and then reused as a gym the agent could enter at will. That promise sat beside a clear historical caution: classical model-based control had always been limited by model misspecification, and learned models inherited the same problem in a less interpretable form. Early deep model-based work therefore mattered less as a finished recipe than as a reframing of the budget: spend experience on a simulator, then spend computation inside it.
Why it matters now
Learned dynamics remain central wherever real interaction is the bottleneck and a perfect simulator is unavailable. Modern agents still face the same trade-off between imagination depth and model fidelity, now at larger scale and often in latent spaces rather than raw observations. The design questions have not gone away: how long to roll out, when to replan from real state, how to stop a policy from exploiting model flaws, and how to keep the dynamics model honest as the policy’s state distribution shifts. Anyone training open controllers against slow or fragile environments still meets this pattern—outer loop on real data, inner loop in a learned simulator—and still judges success by whether real returns track the imagined ones.
The surprising detail
A dynamics model can score well on next-step prediction and still be a poor teacher. The failure is structural: training loss is usually local, while the use case is open-loop rollouts that recycle predictions as inputs. Error that is negligible for one step becomes a different world a few dozen steps later. Agents trained purely in that drifted world often look excellent in imagination and mediocre or unstable on the real system—the gap itself is diagnostic. That is why many working systems deliberately keep imagination short, re-anchor on true observations, or train the policy against an ensemble of models that disagree outside the data, rather than trusting a single average loss number.
What is disputed
There is no single agreed architecture for learned dynamics in RL. How much planning versus policy learning should rely on imagination, how uncertainty should be represented, and how long rollouts may run before they do more harm than good all remain design choices with evidence that depends on the domain. Treat the pattern as a family of methods, not one settled algorithm.
Remember this
A learned dynamics model is a second, cheaper environment—useful only for as many steps as its predictions stay near the truth.
Test yourself
You train a policy entirely on long imagined rollouts from a one-step dynamics model that has low average prediction error on the replay buffer. On the real environment the policy collapses. Give two distinct mechanisms that can produce this gap, and one concrete change to the training loop that addresses each.
First, compounding model error: small one-step mistakes accumulate so late imagined states leave the training support and the policy learns to exploit fantasy transitions. Shorten the imagination horizon, re-anchor rollouts on real states, or stop rolling out when an ensemble’s disagreement spikes. Second, distribution shift: the policy seeks states the buffer rarely contains, so the model was never constrained there. Interleave more real exploration under the current policy, refresh the model on that data, and optionally penalise or constrain actions that drive ensemble variance high. Low one-step loss on old data does not certify multi-step truth under a new policy.
Go deeper
- [1905.08196] Optimisation of Overparametrized Sum-Product Networks · arxiv.org
- [2006.16723] Neural Datalog Through Time: Informed Temporal Modeling via Logical Specification · arxiv.org
Image: Original diagram, The Daily Triptych. Licence: Original work. Source.