Robotics & Embodied AI
A robot policy maps what it currently sees, plus what it's told to do, into what it does next. Almost the whole discipline is one argument about how much of that mapping you hand-engineer with classical control, learn from demonstrations, learn through trial and error, or inherit from a giant pretrained vision-language model.
A robot's software runs perception, state estimation, planning, and control as a pipeline. Learning has crept into every stage of it without replacing the stage below. Classical control (PID, LQR, MPC) still runs the low-level stabilization loop on almost every real robot, learned or not.
It's fast, predictable, and doesn't need training data to work correctly. Above that, imitation learning trains a policy to copy expert demonstrations directly. But plain behavioral cloning drifts off-distribution the moment it makes a small mistake, since nothing in the demonstrations shows how to recover.
Diffusion policies fix the sharper version of this problem by modeling the full, sometimes multimodal, distribution over future action sequences instead of regressing to one averaged answer, and vision-language-action (VLA) models push this further still.
Fine-tune a vision-language model to emit robot actions as just another token type, and web-scale visual and language knowledge transfers straight into physical control. One sentence for an interview: a robot policy is a learned function from observation and instruction to action, and the entire difficulty of robotics is that, unlike text or images, you cannot download more of the real world.
01Intuition
Strip away every acronym and a robot policy is just a function: observation and instruction in, action out.
Two things go in:
- What the robot currently sees. A camera image, a depth map, joint encoder readings.
- What it's being asked to do. An instruction, a goal, or sometimes nothing explicit at all.
One thing comes out: a specific action. Motor commands, a target pose, a change in gripper position.
Every topic on this page answers the same underlying question: how do you build that function, and how much of it should be hand-derived physics instead of learned from data?
A language model can read a trillion tokens scraped from the internet in a few weeks of pretraining.
A robot generating its own training data has to physically move its own arm, in real time, one demonstration or one trial-and-error attempt at a time. Every failed attempt can mean a dropped or broken object, a damaged robot, or a dangerous outcome if a human is nearby.
The real world is far messier, slower, and more expensive to collect data from than any training set of scraped images or text ever is. That single constraint is the root cause behind almost everything else in this section:
- Sim-to-real exists because simulated data is fast, cheap and safe, which real data is not.
- Imitation learning exists because trial-and-error directly on real hardware is slow and can be destructive.
- Hierarchical planning exists to avoid one giant learned function making every decision, from millisecond-level motor torques up to "what task should I even be doing".
Learn only where learning earns its keep. Lean on fast, verified classical control everywhere it doesn't need to be reinvented.
A robot arm's joints have angles, but the task is described by where the tip should go. Going from angles to tip position is called forward kinematics.
Try it Compute where a two-link robot arm ends up
import numpy as np
l1, l2 = 1.0, 0.7 # link lengths in metres
def tip(a1_deg, a2_deg):
a1, a2 = np.radians(a1_deg), np.radians(a1_deg + a2_deg)
return (l1 * np.cos(a1) + l2 * np.cos(a2),
l1 * np.sin(a1) + l2 * np.sin(a2))
for a1, a2 in ((0, 0), (90, 0), (45, 45), (45, -45)):
x, y = tip(a1, a2)
print(f"angles ({a1:>3}, {a2:>3}) -> tip at ({x:.2f}, {y:.2f}) m")
(1.70, 0.00) when straight out and (0.00, 1.70) when raised 90 degrees. At (45, 45) the tip is at (0.71, 1.41), and at (45, -45) it is at (1.41, 0.71). Every set of angles gives exactly one tip position. The reverse question, which angles reach a given point, can have several answers or none, which is why control is harder than this.02Timeline
Robots were controlled by classical control theory (PID, LQR, MPC) with hand-engineered perception and known CAD models, extremely reliable for narrow factory tasks and brittle once an object, lighting or instruction fell outside what was programmed.
Deep RL and imitation learning (2015 onward) learned policies from demonstrations or trial and error, Diffusion Policy (2023) represented multimodal action distributions better than regression, and RT-2-style vision-language-action models (2023) fine-tuned a vision-language model to output robot actions directly.
The frontier is robot foundation models, one policy across many embodiments, with world models (22) increasingly used as learned simulators for planning and for training policies more safely than trial and error on hardware.
03The robot learning stack, piece by piece
Any robot's software can be described as four stages running in a tight loop, fast enough to keep up with a moving, physical world.
- Perception. Turns raw sensor data (camera images, depth, joint encoders, force/torque readings) into something usable: object positions, a segmentation mask, a point cloud.
- State estimation. Fuses that noisy, partial sensor stream over time into a best current estimate of what's true. Classically a Kalman filter or one of its nonlinear variants: combine what the sensors say with what the dynamics model predicts, weighted by how much you trust each.
- Planning. Decides what to do next to make progress toward a goal, from a high-level "pick up the mug" down to a concrete trajectory.
- Control. Turns that plan into actual motor commands, moment to moment, correcting for disturbances as they happen.
Deep learning has steadily taken over pieces of this pipeline. Perception first, then increasingly planning and parts of control. But not all of it. Knowing which parts are still classical, and why, is most of what separates real understanding here from hand-waving.
Classical control still runs the last mile: PID, LQR, MPC
Three controllers, in increasing order of how much they assume about the system.
- PID (proportional-integral-derivative) is the simplest and by far the most common. It reacts proportionally to the current error, meaning how far you are from the target. It adds a term proportional to the accumulated (integral) error over time, which kills steady-state offset. And it adds a term proportional to how fast the error is changing (derivative), which damps oscillation before it happens. Three gains, no dynamics model required, simple enough to run on the microcontroller inside a single motor.
- LQR (linear-quadratic regulator) goes a step further. Write down a linear model of the system's dynamics and a quadratic cost function that penalizes deviation from the target and penalizes control effort. LQR then gives you the mathematically optimal linear feedback controller for that model, in closed form. No search, and provably optimal for the model you handed it.
- MPC (model predictive control) goes further still. At every timestep it uses a model of the system to plan a short trajectory a few steps into the future, minimizing a cost subject to constraints like joint limits and obstacle avoidance. It executes only the first action of that plan, throws the rest away, and re-plans from scratch at the next timestep with fresh sensor data.
All three are decades old, unglamorous, and still doing an enormous amount of the actual work in modern robots. That includes robots with a state-of-the-art learned perception system and a learned high-level policy sitting on top.
The loop that keeps a robot arm from swinging wildly, or a legged robot from faceplanting, has to be fast (kilohertz-range control rates), provably stable within its operating envelope, and completely predictable. A learned policy doesn't inherently have those properties, and would need substantial extra engineering to guarantee them.
Learning earns its keep at the level of what should the robot try to do. Object recognition, grasp selection, task sequencing. Not necessarily at the level of exactly how many newton-meters of torque, right now, to keep this joint where I told it to be.
Inverse kinematics and trajectory optimization
Forward kinematics is the easy direction: given a robot's joint angles, you compute where its end-effector ends up in space with chained geometry down the arm.
Inverse kinematics (IK) asks the harder, more useful question backwards. Given a desired end-effector pose, what joint angles achieve it? For arms with more joints than the six degrees of freedom needed to reach an arbitrary pose, there are often infinitely many valid solutions. An IK solver has to pick one, typically by optimizing for something like minimal joint movement or staying away from joint limits.
Trajectory optimization extends IK over time. It solves for a whole sequence of configurations, in place of one static joint configuration, that moves the arm from where it is to where it needs to be: smoothly, without collisions, while respecting velocity and acceleration limits. This is a constrained optimization problem that is solved once or, in the MPC case above, continuously re-solved.
Sim-to-real and domain randomization
Real robots produce data at the speed of the real world, no faster. One demonstration takes exactly as long as physically performing the task. And every attempt risks damaging the robot, the object, or, with a person nearby, someone's safety.
Simulation removes all three constraints. Thousands of instances can run in parallel, faster than real time, on a GPU cluster. A simulated robot that slams its gripper into a wall costs nothing but compute.
The catch is the sim-to-real gap: a policy trained against a simulator's specific rendering, a specific friction model, and slightly wrong mass and inertia values for every object will quietly learn to rely on those exact, wrong details. It then fails in the real world in ways that look arbitrary from the outside, because it was solving the specific simulated instance and never the general task.
Domain randomization is the mechanistic fix, and it is not just "add some noise". During training, aggressively randomize exactly the simulation parameters the real world is uncertain about:
- Textures and lighting
- Camera pose and intrinsics
- Object mass and friction coefficients
- Motor latency and sensor noise
Randomize across a range wide enough that the policy is forced to become invariant to all of them, and cannot exploit any one specific value. Done aggressively enough, the real world stops looking like a novel, out-of-distribution input and looks like one more randomized simulation, sampled from the same wide distribution the policy already trained across.
Imitation learning and behavioral cloning: the compounding error problem
Imitation learning trains a policy by direct supervised learning. Collect observation-action pairs from an expert, usually a human teleoperating the robot, then train the policy to reproduce the expert's action given the observation.
This is behavioral cloning, and it is appealing. There is no reward function to design and no exploration problem: it is ordinary supervised learning on a dataset you already know how to collect.
Its weakness is compounding error. Expert demonstrations only show near-perfect trajectories, so the trained policy will make small mistakes the expert never made, and each mistake nudges the robot into a state slightly outside anything the demonstrations covered. The policy was never trained on that state, so its next prediction there is worse, and the next mistake compounds on top of the last one.
Compare that to a misclassified held-out image. A bad action here changes the world the policy has to act in next. And nothing in a dataset of only successful, on-distribution demonstrations tells the policy how to recover once it's off-distribution.
One direct, well-known mitigation is DAgger (dataset aggregation). Instead of training purely on offline expert demonstrations, it runs the current policy and has the expert label the states it visits, including its mistakes. Those corrections go back into the training set, which closes the distribution-shift gap by training on the policy's own error states and does not depend on the original demonstrations happening to cover them.
Action chunking helps in a different, complementary way. Instead of predicting one action at a time, the policy predicts a short sequence of several future actions at once, conditioned on the current observation.
That reduces how many independent decision points there are per unit of task progress. A whole chunk is generated jointly and is internally consistent, where single-step guesses are stitched together and could each drift independently. Fewer decision points leave fewer places for compounding error to start.
Diffusion policies: actions as the thing being denoised
Diffusion Policy (Chi et al., 2023) takes the diffusion machinery from 12 and points it at a different target. Instead of denoising an image, the network denoises a sequence of future robot actions, conditioned on the current observation.
Sample pure noise over an action chunk. Run the same iterative learned-denoising process covered in 12. What comes out is a plausible, coherent action sequence given what the robot currently sees. Same math, a different object being generated.
Beyond being the newer diffusion technique, the reason this fits robot control specifically is that it naturally represents multimodal action distributions in a way a single regression output structurally cannot.
A mug's handle can be grasped equally well from the left or the right. Both show up in the demonstrations.
A plain regression policy trained to minimize squared error against both doesn't pick one and averages them instead, which gives an output that is physically invalid at both extremes and unhelpful in between.
A diffusion policy can place probability mass on the left-grasp mode and the right-grasp mode separately, then sample a coherent one. It's the same property that lets the image diffusion models in 12, which can produce several different plausible images from the same prompt, avoid blending them into something no expert demonstration ever did.
Vision-language-action (VLA) models: actions as tokens
A VLA model takes in an image and a natural-language instruction and outputs robot actions directly. You build one by taking an already-pretrained vision-language model and fine-tuning it to also emit action tokens, alongside the text tokens it already knew how to produce.
RT-2 (Brohan et al., 2023, Google DeepMind) is the model that made this idea concrete. Each dimension of a continuous robot action gets discretized into bins: small changes in end-effector position, orientation, gripper openness.
Each bin is represented as a token drawn from the model's existing vocabulary. The model is then fine-tuned on a mix of web-scale vision-language data and robot trajectory data, predicting those action tokens the same way it predicts the next word in a sentence.
This is the everything-is-tokens idea from 10, pushed one step further: the vocabulary is no longer only sub-words and also has entries that mean "move the gripper slightly left."
The payoff is large. Broad visual and semantic knowledge learned from internet-scale image-text data transfers directly into robot control: what a knife looks like, what "the smallest object" means, common object affordances. That lets the model follow instructions and recognize objects far outside anything a robot-specific dataset alone could have taught it.
Robot foundation models and cross-embodiment learning
A model trained on one robot's data doesn't automatically work on a different robot. Its specific arm geometry, gripper, and camera placement are baked in, and the new robot has a different action space and different dynamics.
Cross-embodiment learning trains a single policy across data pooled from many different robot bodies at once, where the alternative is a bespoke policy per robot. The bet is that a policy exposed to enough embodiment diversity learns something closer to a general, embodiment-agnostic notion of "how to manipulate objects" and does not memorize one specific robot's quirks.
It mirrors exactly what made large language and vision models generalize well: scale and diversity of training data beating a narrow, hand-curated dataset. Here that idea is applied to the harder problem of physical bodies that don't even share an action space.
This remains an active, unresolved research direction. Getting one policy to transfer skill across a wheeled arm, a legged robot, and a humanoid hand is a much harder generalization problem than transferring across different image resolutions.
Language-conditioned manipulation, navigation, and mobile manipulation
Language-conditioned manipulation means the policy takes a natural-language instruction as an additional input alongside the visual observation. The same underlying model then performs different tasks depending on what it's told, without a separately trained policy per task.
Navigation is the distinct problem of moving the robot's base through space. Build or use a map (classically via SLAM, simultaneous localization and mapping), plan a collision-free path, and follow it.
Mobile manipulation combines the two, and the coordination between them is the hard part. The robot has to navigate to a position from which the target object is reachable and visible, and "close enough" is not good enough, so the navigation planner and the manipulation planner can't be designed in isolation from each other.
Humanoid learning and tactile sensing
Humanoid robots add a constraint most manipulators don't have. The robot has to balance itself while doing everything else.
Locomotion (walking, recovering from a push) is typically handled by an RL-trained or classical whole-body controller running underneath, enforcing balance constraints. Manipulation and higher-level task policies sit on top of it. That's another instance of the same hierarchical pattern: fast safety-critical control at the bottom, learned flexible decision-making above it.
Tactile sensing fills a gap vision structurally can't close. A camera can't see what's happening at the point of contact once a hand occludes it. And vision alone often can't distinguish a firm grip from a slipping one until the object is already falling.
Tactile sensors measure contact force, pressure distribution, and slip directly at the fingertip. Some designs turn touch into an image, using a camera behind a soft deformable membrane, which lets the same vision architectures used elsewhere in the stack process touch data too.
World models as learned simulators, and hierarchical planning
A world model, in the robotics sense, is a learned function that predicts what happens next. Given the current observation or state and an action, it predicts the next one, without needing to execute that action in the real world or a hand-built simulator.
Once you have that, you can plan by imagining rollouts inside it. Try several candidate action sequences, see which one the model predicts leads somewhere good, and execute only that one for real. No purely real or simulated trial and error required.
22 covers this in depth. The short version here is that a good learned world model turns planning into a search problem you can run cheaply and safely inside the model's imagination first.
This connects back to the stack diagrammed above. Three different systems operate at three different timescales and cooperate, where the alternative is one giant end-to-end model making every decision from "what task to do" down to "how many newton-meters, right now."
- A language or task-level planner picks a sub-goal a few times a minute.
- A skill policy, which might itself be a diffusion policy or a VLA model, figures out how to achieve that sub-goal, replanning several times a second.
- A low-level controller, often still classical, stabilizes and executes the resulting motion hundreds to thousands of times a second.
Each layer solves a different problem at a different rate. Forcing one model to do all three tends to produce something simultaneously too slow for the fast layer and too myopic for the slow one.
Safety, uncertainty, and recovery
A learned policy that's confidently wrong is more dangerous than one that knows it's uncertain.
Real deployed systems need the policy, or a layer wrapped around it, to recognize when it's operating outside anything it was trained on. It should then fall back to a safe, conservative behavior: stop, ask for help, or hand control back to a human or a classical safety controller. Any of these beats confidently executing an action from a part of the input space it has never seen.
Diffusion policies have a structural advantage here over plain regression. The shape of the predicted distribution, tight and confident versus spread out and uncertain, is itself a signal. A single regression output is just a number, with no attached sense of how much to trust it.
On top of whatever the learned policy does, most real systems still keep a hard, classical safety layer underneath: force and torque limits, joint limit enforcement, emergency-stop conditions. It can override the learned policy's output before it ever reaches the motors. A safety-critical guarantee shouldn't rest entirely on a learned system behaving well.
Open-loop vs. closed-loop evaluation
Open-loop evaluation checks a policy against a fixed, pre-recorded scenario. Replay a logged sequence of observations and check whether the policy's predicted actions match the expert's logged actions at each step.
It is cheap, fast and reproducible, but it can't catch compounding error, because the robot's own actions never change what it sees next. Every input it is evaluated on was generated by the expert and not by the policy.
Closed-loop evaluation runs the policy for real, in simulation or on hardware, and lets its own actions determine what it sees at the next step. That's the same way it will be evaluated in deployment.
Closed-loop evaluation is more informative, and it is the only way to catch compounding-error failures before they show up in the real world. It is also slower, more expensive, and on real hardware riskier to run at scale.
The practical pattern: iterate fast on open-loop metrics and closed-loop simulation, then treat real-hardware closed-loop evaluation as the final, expensive gate before deployment and not the primary development loop.
04The core objective
Plain behavioral cloning trains a policy with ordinary regression: $\mathcal{L}_{BC} = \mathbb{E}_{(o,a)\sim\mathcal{D}}\big[\lVert \pi_\theta(o) - a\rVert^2\big]$. Given the observation $o$, predict the expert's action $a$, and minimize squared error.
The mode-averaging problem from the diffusion policy discussion above shows up mathematically here. If the demonstration data contains two equally valid ways to grasp an object, the single point estimate that minimizes squared error against both lands between them, which is physically valid for neither.
That's a separate, additional failure on top of compounding error, and it's the specific thing diffusion policies fix.
A diffusion policy's training objective is the DDPM loss from 12, with one substitution: the thing being denoised is now a chunk of future actions, conditioned on the current observation, in place of an image.
- a_0 is the ground-truth action chunk, a short window of future actions read off a demonstration and not a single timestep's action.
- o the current observation (image plus proprioception) — the conditioning signal, playing exactly the role a text prompt plays in 12's conditional diffusion.
- t, ε, ᾱ_t identical machinery to 12: a random diffusion timestep, the Gaussian noise added, and the fixed noise schedule. See that page for the full derivation — this is the same objective with "image" swapped for "action sequence."
At sampling time you start from pure noise over the action chunk. Run the learned reverse process conditioned on the current observation, execute some of the resulting actions on the robot, re-observe, and denoise again for the next chunk.
DDIM-style fast samplers (12) apply here too, and they matter more in robotics than in image generation. A policy that takes several seconds to decide on its next action is often unusable at a real control rate.
05Why not just imitate, and why not just collect more real data
Behavioral cloning is supervised learning wearing a robot costume, and it inherits supervised learning's central assumption: train and test distributions match.
That assumption breaks specifically because the robot's own actions determine its next input. A wrong prediction doesn't just cost you one bad label. It moves the robot into a state the training data never covered, and there's no signal anywhere in a dataset of only successful demonstrations telling the policy how to get back.
Three techniques attack that from different angles:
- Diffusion policies model the full action distribution in place of one averaged point estimate, which avoids confidently executing a physically invalid "average" action.
- Action chunking reduces the number of independent decision points per unit of task progress where drift can start.
- RL fine-tuning optimizes directly for task success using the robot's own on-policy experience, including its failure states. That's exactly the recovery signal pure imitation never gets, because a human demonstrator basically never generates a demonstration that starts from "the robot already made a mistake."
The case for simulation over collecting more real data is a matter of what's feasible to scale. Real robot data collection is bounded by four things:
- Real time. A demonstration takes as long as physically doing the task.
- Physical wear accumulates on motors and grippers.
- Safety around fragile objects and people.
- Cost. A fleet of real robots running around the clock is enormous capital expense next to thousands of parallel simulated instances running faster than real time on a shared GPU cluster.
Simulation flips every one of those constraints, at the price of the sim-to-real gap, and domain randomization is the direct answer to that gap. If you randomize exactly the simulation parameters the real world is uncertain about, aggressively and across a wide range, the policy is forced to become invariant to them and cannot exploit any single simulated value.
Push that range wide enough and the real world looks like one more sample from the training distribution and is no longer a novel input.
06Complexity, failure modes, and what breaks
| Classical control | Imitation learning (BC) | RL | VLA models | |
|---|---|---|---|---|
| Data requirement | None — hand-derived model & gains | Hundreds to thousands of demonstrations | Large volume of trial-and-error rollouts, sim or real | Web-scale pretraining + robot demo/rollout data |
| Generalization | None beyond the specified task/model | Weak — only as good as demo coverage | Can exceed demos, but only if the reward is well-specified | Strongest to new objects/instructions — inherits the VLM's priors |
| Safety guarantees | Strong, provable stability margins for known/linear systems | None inherent — as unpredictable as the demos it copied | None inherent — can find high-reward but unintended behavior | None inherent — also inherits the VLM's confidently-wrong tendencies |
What shows up in production
- Compounding error appears in naive behavioral cloning. Small mistakes push the robot off the demonstrated distribution, and there's no recovery signal in the training data to pull it back.
- Sim-to-real gap when domain randomization isn't aggressive or wide enough. The policy quietly relies on some specific simulated detail (a friction value, a camera intrinsic) that doesn't hold in reality, then fails in a way that looks arbitrary from the outside.
- Inherited hallucination in VLA models, now applied to physical actions. A language model hallucinating a fact is embarrassing. A VLA model misjudging whether an object is graspable, or misreading which object an ambiguous instruction refers to, moves a real gripper into a real, possibly damaging, action. Higher stakes than text hallucination.
- Reward hacking appears in RL, where the policy optimizes what you specified and not what you meant. A degenerate but technically-rewarded motion found on physical hardware is a safety issue, however amusing it reads in a paper.
"It was fine-tuned from a strong vision-language model, so it must understand physics." Not necessarily. It inherited strong visual recognition and instruction-following, so it's good at telling you what's in the scene and what you asked for.
Physical competence is a different thing. Whether a grasp will hold, whether an object is too heavy, whether a plan is dynamically feasible. That has to come from the action-prediction fine-tuning on real robot interaction data, and that data is orders of magnitude smaller than the web-scale pretraining corpus.
Treat "built on a strong VLM" as a claim about perception and instruction-following, which gives no guarantee of physical grounding.
Diffusion policies and VLA models are the leading research directions, and are increasingly used in early production for structured tasks like warehouse pick-and-place. But they haven't replaced classical control at the low level. Nearly every deployed robot, however modern its perception or planning stack, still stabilizes its own joints with PID, LQR, or MPC running underneath.
Robust cross-embodiment generalization, meaning one policy performing well across different robot bodies without extra fine-tuning, remains an open research problem. Treat strong claims of a single "generalist robot policy" working out of the box on a brand-new robot body as an overclaim unless backed by specific, verifiable evaluation.
07Build this
Action chunking reads like a batching trick. Plot the actions a chunked policy executes next to a single-step one, and it stops looking like plumbing.
Clone the same expert twice on a simulated reach task. One policy predicts the next action. The other predicts a chunk of the next several. The two share data, loss and network and differ in one way. A reach suits this because there is exactly one sensible motion, so anything ugly in the executed trajectory came from the policy and the task is not to blame.
- Pick a reach task with a low-dimensional state, such as Gymnasium's
Reacher. Write the expert as a simple PD controller driving the tip toward the target, then record a few hundred episodes of(observation, action). That expert is the classical control layer from section 03, and it costs you nothing to write. - Train policy A straight from the behavioural-cloning loss in section 04. An MLP maps the observation to the single next action, squared error against the expert's.
- Train policy B. It uses the same MLP, data and loss, except that the output layer emits the next $K$ actions at once. Read the targets off each demonstration as a sliding window. Try $K$ around 8 or 16.
- Roll both out closed-loop. A acts once per observation, and B executes its whole chunk, then re-observes and predicts the next one. Record every action that reached the simulator.
- Plot the executed action for each joint against time, A and B on the same axes. Plot the two tip paths beside them.
- Now break it deliberately by rewriting the expert so it reaches the target two different ways at random, swinging clockwise or anticlockwise around the midpoint. Retrain both policies on that data and plot again.
Where this runs in production
A common real deployment pattern runs in three layers:
- A learned perception-and-policy module looks at a bin of mixed, previously-unseen items and outputs where and how to grasp: position, approach angle, gripper width. That's a high-level decision that needs generalization.
- A classical trajectory-optimization and IK layer computes the joint-space path to reach the chosen grasp.
- A low-level PID or MPC controller executes that path and stabilizes the arm against disturbances in real time.
The learned part handles what requires generalization. Classical control handles what doesn't need to be learned at all, and benefits from being fast, predictable, and easy to verify.
This split, learn the high-level decision and keep the low-level execution classical, is the dominant pattern in deployed robotics right now. That split is a durable design, and it explains why the tradeoffs table above has no single winner.
08Interview questions
BeginnerWalk me through the perception → state estimation → planning → control pipeline.
Perception turns raw sensor data — camera images, depth, joint encoders — into something usable, like object positions or a point cloud. State estimation fuses that noisy, partial data over time into a best current estimate of the true state, classically with a Kalman filter or a nonlinear variant. Planning decides what to do next to reach a goal, from a high-level instruction down to a concrete trajectory. Control turns that plan into actual motor commands, correcting for disturbances moment to moment. Learning has replaced pieces of each stage over time, but not uniformly — perception first, planning and control more gradually.
BeginnerWhy do modern robots still use PID controllers when we have deep learning?
Because the low-level stabilization loop needs to be fast (kilohertz-range), provably stable within its operating envelope, and completely predictable — properties a learned policy doesn't inherently have. PID, LQR, and MPC give you those guarantees directly from a known model, with no training data required. Learning earns its keep at the level of deciding what the robot should try to do, not necessarily at the level of exactly how much torque to apply right now to hold a joint in place.
IntermediateWhy does behavioral cloning suffer from compounding error, and how do you mitigate it?
Behavioral cloning trains on demonstrations that are near-perfect, so the training distribution only contains on-track states. The learned policy will still make small mistakes, and each one nudges the robot into a state slightly outside the demonstrated distribution — a state the policy was never trained on, so its next prediction there is worse, and the error compounds. Nothing in a dataset of only successful demonstrations shows how to recover. Mitigations: DAgger, which queries the expert for corrective labels on states the current policy actually visits, closing the distribution-shift gap directly; action chunking, which predicts several actions jointly per observation instead of one at a time, reducing the number of independent points where drift can start; and RL fine-tuning, which gives the policy a direct success signal from its own on-policy experience, including failure states imitation data never contains.
IntermediateWhat does domain randomization actually do, mechanistically?
It randomizes simulation parameters the real world is uncertain about — textures, lighting, camera pose and intrinsics, object mass, friction, motor latency, sensor noise — aggressively and across a wide range during training. That forces the policy to become invariant to those parameters instead of overfitting to one specific simulated value. If the randomization range is wide enough, the real world doesn't look like a novel, out-of-distribution input to the policy — it just looks like one more randomized simulation instance, sampled from the same broad training distribution.
IntermediateWhy are diffusion policies better suited to robot control than a plain regression policy head?
A plain regression head trained with MSE against multimodal demonstration data — say, two equally valid ways to grasp an object — doesn't pick a mode, it averages them, producing an output that's physically invalid at both extremes and unhelpful in between. A diffusion policy models the full distribution over action sequences and can sample from one mode coherently instead of blending across modes. It's the same reason image diffusion models can produce several genuinely different valid images from one prompt rather than one blurry average.
DeepHow would you evaluate a robot policy before deploying it on real hardware?
Start with open-loop evaluation against logged expert trajectories — cheap and fast, but it can't catch compounding error, since every input it's tested on came from the expert, not from the policy's own actions. Move to closed-loop evaluation in simulation, where the policy's own actions actually determine its next observation, which is the only way to expose distribution-shift failures before they happen on hardware. Only after that would you run closed-loop evaluation on real hardware, treated as the final, expensive gate rather than the primary development loop, ideally with a hard classical safety layer — force limits, joint limits, an e-stop condition — that can override the policy regardless of what it outputs.
DeepHow does a VLA model like RT-2 turn a language model into something that outputs robot actions?
By discretizing each continuous dimension of the robot's action — small changes in end-effector position, orientation, gripper openness — into bins, and representing each bin as a token drawn from the model's existing vocabulary. The pretrained vision-language model is then fine-tuned on a mix of web-scale vision-language data and robot trajectory data to predict those action tokens the same way it predicts the next word in a sentence — actions become just another entry in the output vocabulary, the same everything-is-tokens framing used for text. That's also why VLA models transfer broad visual and semantic knowledge from web pretraining directly into robot control, and why they can inherit a language model's confidently-wrong failure modes, now applied to a physical action instead of a sentence.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.