←Home KnowML
Embodied & FrontierChapter 20

3D, Spatial AI & Autonomous Driving

Every sensor on a self-driving car gives you an incomplete, differently-shaped view of the world. Cameras throw away depth, LiDAR throws away color and density, and radar throws away almost everything except velocity. This page is about the geometry and the learned representations that stitch those partial views into one 3D picture accurate enough to bet a car's braking decision on.

24 min read Assumes: CNNs & detection basics (05), attention & transformers (08)
Start reading
TL;DR

Camera geometry (intrinsics, extrinsics, epipolar constraints) is the mathematical glue. It lets you take pixels from several differently-positioned cameras and project them into one shared 3D frame, or more commonly a bird's-eye-view (BEV) frame. Once everything lives in that shared frame, detection, tracking, prediction and planning all become standard learned problems on top of it, rather than per-camera special cases.

Modern perception stacks fuse camera, LiDAR and radar because each sensor fails differently. Cameras are cheap and semantically rich but bad at depth and useless in the dark. LiDAR gives precise 3D geometry but is expensive and sparse at range. Radar is the only one measuring velocity directly via Doppler, at very coarse resolution.

BEV transformers (BEVFormer) and occupancy networks (TPVFormer) are the current best answer for fusing them into one representation, instead of stapling together separate per-sensor detectors. Once you have that fused representation, trajectory prediction forecasts where every other agent is going next, and motion planning turns that into a safe, comfortable trajectory for your own car.

Almost every deployed system keeps these as separate, debuggable stages instead of one end-to-end network, because "why did the car do that" needs a debuggable answer as well as a working one.

One sentence for an interview: autonomous driving is a real-time 3D perception problem wearing a robotics costume, and its central engineering tension is fusing complementary, imperfect sensors into one consistent world model fast enough, and reliably enough, to plan around safely.

01Intuition

No single sensor tells you what's out there, so the whole field is about combining several incomplete answers into one you can trust.

A camera image is a 2D projection of the world. By the time light hits the sensor, depth is already gone. Two objects at wildly different distances can land on the exact same pixel if one sits behind the other, and a monocular camera alone cannot tell you which is which. It needs additional assumptions, or a second view to triangulate against.

The other two sensors each fix part of that, and break somewhere else:

  • LiDAR fires laser pulses and times their return, giving precise, direct 3D range measurements. But a rotating unit only samples a sparse, uneven point cloud: dense up close, sparse at range, and silent about color or fine texture. It is also the most expensive sensor in the stack.
  • Radar measures velocity directly and instantly via Doppler shift, which neither cameras nor LiDAR do without inferring it across multiple frames. Its spatial resolution is coarse enough that it often cannot tell you an object's exact shape, or sometimes whether two nearby objects are one object or two.
The two questions everything here answers

First: how do you turn what one sensor sees into a 3D fact about the world? That is camera geometry, stereo depth, and LiDAR point-cloud processing.

Second: how do you combine facts from sensors that disagree? They measure different things and update at different rates. Reconciling them is sensor fusion, BEV representations, and occupancy networks.

Get the first right and you have accurate individual readings that still don't agree with each other. Get the second right and that disagreement flips from a source of confusion into a source of robustness, because each sensor covers the others' blind spots.

A camera turns a point in 3D into a pixel by dividing by its distance. That one division is why far things look small and why a single camera can't tell you how far away something is.

Try it Project a 3D point onto a camera image
f = 800.0                                    # focal length in pixels
cx, cy = 320.0, 240.0                        # image center

def project(X, Y, Z):
    return f * X / Z + cx, f * Y / Z + cy    # x = f X / Z

for Z in (5, 10, 20):
    u, v = project(1.0, 0.5, Z)
    print(f"{Z:>2} m away -> pixel ({u:.0f}, {v:.0f})")
The same point, 1 m right and 0.5 m down, lands at pixel (480, 320) at 5 m, (400, 280) at 10 m and (360, 260) at 20 m. Moving it further away pulls it toward the image center at (320, 240). Going the other way, from a pixel back to a 3D position, needs the distance, and one image doesn't contain it.

02Timeline

Before (2000s – early 2010s)

Depth came from structure-from-motion and classical stereo, tracking from SIFT-style features and Kalman filters, and planning from explicit rules, all rigorous and interpretable but each piece degraded in glare, rain and unanticipated object shapes.

→
Innovation (2016 – 2022)

PointPillars (2018) made LiDAR detection fast and learnable, ORB-SLAM3 (2020) fused camera and IMU in one optimization, and BEVFormer (2022) lifted several camera views into one bird's-eye-view grid with learned cross-attention, so features are fused first and detection happens once.

→
After (2023 – 2026)

Occupancy networks predict free space and semantics per voxel, so odd obstacles still show up as occupied, NeRF and Gaussian splatting (06) reconstruct driving logs for simulation, and modular-but-learned versus end-to-end is still unsettled (section 05).

03The stack, piece by piece

A modern autonomous driving stack is a pipeline, much like the robotics pipeline in 21, but with 3D geometry doing more of the early work. Raw sensors feed geometric and learned perception. Perception output gets fused into one shared representation. That representation feeds prediction and planning. Planning's output becomes control commands.

Sensors → shared 3D representation → plan — one pipeline, several sensor streams
sensors Cameras (×N) LiDAR Radar Fusion: BEV grid / occupancy volume one shared 3D frame Detection & tracking Trajectory prediction Motion planning Control (21's PID/MPC)
Three sensor streams, running at different rates and with different failure modes, get lifted into one shared bird's-eye-view or occupancy representation before anything downstream ever runs. Detection and tracking, prediction, and planning are almost always kept as separate stages — see section 05 for why.

Camera geometry: intrinsics, extrinsics, and why multi-camera needs both

Two transforms, doing two different jobs:

  • The intrinsic matrix $K$ describes properties fixed by the lens and sensor: focal length, and the principal point where the optical axis hits the sensor. It maps a 3D point already expressed in the camera's own frame down to a 2D pixel.
  • The extrinsic transform, a rotation $R$ and translation $t$, describes where that camera physically sits and points relative to the vehicle. It converts a point from world coordinates into that camera's frame, before the intrinsics project it to a pixel.

A single camera only ever needs its own intrinsics and extrinsics. A multi-camera rig needs every camera's extrinsics relative to one shared reference frame, typically the vehicle's own center. That shared frame is what lets you take features from six differently-mounted cameras and place them correctly relative to each other, instead of holding six disconnected pixel grids.

Epipolar geometry, stereo, and depth

Given two cameras with a known, fixed offset, the same real-world point appears in both images. Epipolar geometry says something powerful about where. Once you know the two cameras' relative pose, a point's location in the first image constrains its possible location in the second to a single line, the epipolar line, rather than the whole 2D image.

That collapses stereo matching from a 2D search into a 1D one. Which is the difference between "compute this on every frame in real time" and "don't."

Once a correspondence is found, depth falls out of the geometry directly. Farther points shift less between the two images (smaller disparity), and nearer points shift more. For a fixed baseline, disparity is inversely proportional to depth.

SLAM and visual-inertial odometry

Simultaneous localization and mapping is a genuine chicken-and-egg problem. To know where you are, it helps to already have a map. To build a map, it helps to already know where you are. SLAM solves both at once, jointly estimating the camera's trajectory and a map of tracked landmarks, refining both as new observations arrive.

Visual-inertial SLAM tightly couples camera tracking with an IMU's acceleration and angular velocity readings. ORB-SLAM3 is the well-known open-source reference implementation. The two sensors cover for each other:

  • The IMU fills in fast, high-frequency motion between camera frames, and stays useful through the camera's worst moments: motion blur, brief occlusion, a tunnel.
  • The camera corrects the IMU's slow drift, whenever it can see stable landmarks again.

3D object detection: camera-based and LiDAR-based

LiDAR-based detection has to solve an awkward data-structure problem first. A point cloud is sparse, unordered, and wildly uneven in density: dense close to the sensor, sparse far away. A standard CNN, built for dense regular pixel grids, cannot consume that directly.

PointPillars answers it in three moves:

  • Discretize the ground plane into a grid of vertical columns, the "pillars."
  • Run a small shared network over the points inside each pillar, producing one feature vector per pillar.
  • Arrange those vectors into a 2D pseudo-image that an entirely ordinary 2D CNN detector can process.

That turns an irregular 3D problem into a regular 2D one, with no 3D convolutions at all. Which is exactly why it runs fast enough for a moving vehicle.

Camera-based 3D detection has the harder job: estimating depth that was never directly measured. It can do that per-camera, via monocular depth estimation, which is generally less accurate. Or jointly across camera views, the way the BEV methods below do. That accuracy gap is a large part of why multi-camera BEV approaches displaced monocular 3D detection as the camera-only default.

BEV representations: one shared top-down frame

Bird's-eye-view is a top-down 2D grid over the ground plane around the vehicle. It is the natural shared coordinate frame for driving specifically, because planning a route and reasoning about lanes is inherently a top-down problem. Nobody plans a maneuver by thinking in raw camera-pixel coordinates.

BEVFormer builds this grid directly with attention. A fixed grid of learned BEV queries, one per top-down cell, uses two attention mechanisms:

  • Spatial cross-attention pulls in the relevant image features from whichever camera views see that patch of ground.
  • Temporal self-attention fuses the current BEV grid with the previous timestep's grid.

That temporal half matters more than it first looks. It noticeably helps estimate a moving object's velocity, and recovers objects briefly occluded a frame ago, because the model is not restarting perception from zero every single frame.

Occupancy networks: past the bounding box

A bounding box assumes the world is made of a fixed, known set of object categories with roughly box-shaped extents. That is a bad assumption for a fallen ladder, a spilled load of cargo, or any obstacle shape a detector's training set never named.

Occupancy networks sidestep it: for every cell in a grid around the vehicle they predict whether it is occupied and, if so, by roughly what kind of thing. That turns "list the objects" into "describe free space directly", which degrades far more gracefully on anything unusual.

TPVFormer's trick for making this tractable: represent the volume as three orthogonal 2D planes, the top-down view plus the two perpendicular side views, instead of one dense 3D voxel grid. Combining the three planes recovers the full 3D structure at meaningfully lower cost than predicting every voxel directly.

NeRF, Gaussian splatting, and scene reconstruction

06 covers how NeRF and Gaussian splatting work as general 3D scene representations. The driving-specific use is reconstruction for simulation.

Take a real recorded driving log. Reconstruct the actual scene the car drove through as a NeRF or a set of Gaussians. Then re-render it from camera viewpoints the car never took, or with other actors inserted or removed. That gets much closer to testing against reality than a purely synthetic simulator, because the environment was reconstructed from a real drive and not hand-modeled.

Scene graphs and spatial reasoning

A scene graph represents the world as nodes (vehicles, pedestrians, lanes, traffic signals) connected by labeled relational edges: "in front of," "yielding to," "occupies this lane." That is richer than a flat, unstructured list of detections, and it makes higher-level spatial queries answerable directly. "Which vehicles are relevant to my current lane change" becomes a graph query over relationships, rather than a re-derivation from raw geometry every time it is asked.

Sensor fusion: camera, LiDAR, and radar together

Fusion can happen at three points in the pipeline, and where you put it is the whole design decision:

  • Early. Combine near-raw sensor data before much processing. Hardest to get right, but no information is thrown away before combination.
  • Late. Run separate full detectors per sensor and combine only their final outputs. Simplest to build and debug, but each detector may have already discarded a cue only visible when signals are considered jointly.
  • In between. Fuse intermediate features. BEVFormer-style camera fusion, extended to also ingest LiDAR or radar features into the same shared grid, is the currently favored middle ground.
Why these three sensors specifically

Camera, LiDAR and radar are a deliberately complementary trio, and "more sensors are better" does not describe them.

Camera supplies rich semantic detail that the other two structurally cannot: what is this, what color, what does the sign say. LiDAR supplies precise geometry that monocular camera depth estimation cannot match. Radar supplies direct, instantaneous velocity, and keeps working through fog, rain and glare that badly degrade cameras and, to a lesser extent, LiDAR.

Each one's weakness is another's strength.

Trajectory prediction and motion planning

Trajectory prediction takes each tracked agent's recent motion history plus map context (lanes, right-of-way, traffic controls) and forecasts several plausible future paths. Several paths and not one, because "what will that pedestrian do" is inherently multimodal. They might cross or they might not, and a model that averages both possibilities into a single path is wrong in exactly the way 21's behavioral-cloning discussion describes.

MotionNet illustrates one clean way to structure this. It takes BEV maps as input and predicts, per grid cell, both a category and a short-horizon motion vector. Detection and near-term prediction live inside one shared BEV representation instead of two disconnected models.

Motion planning then takes the ego vehicle's goal plus every predicted agent trajectory, and searches for a path that is:

  • Safe. Does not intersect anyone else's likely future path.
  • Comfortable. Bounded acceleration and jerk.
  • Legal. Obeys lane and traffic-control constraints.

Conceptually this is the same constrained trajectory-optimization idea as MPC in 21. The difference is scope: one long route rather than one short re-planned step, re-run continuously as new predictions arrive.

End-to-end autonomous driving, and world models for simulation

End-to-end driving trains one network to map sensor inputs directly to a planned trajectory, or even to control outputs. It optimizes jointly against one end-task loss instead of stage-by-stage intermediate losses.

In principle that lets it discover something the modular pipeline structurally cannot see: that a slightly different detection threshold would have produced a much better downstream plan. Section 05 covers the tradeoffs, and why the industry has not broadly converged on this for safety-critical deployment.

World models (22) increasingly get used here as fast, learned simulators. Instead of only testing planning against a fixed replay log or a hand-built physics simulator, you roll candidate plans forward inside a learned model of how the scene evolves. Cheaply, and safely, before anything touches the real world.

Closed-loop evaluation for driving

The same open-loop-versus-closed-loop distinction from 21 applies here, with higher stakes.

  • Open-loop evaluation replays a fixed recorded log and checks whether the planner's outputs match what happened. It is cheap, fast and fully reproducible, and also blind to compounding error, because the ego vehicle's hypothetical actions never change what the scenario does next. A bad early decision never gets the chance to cascade.
  • Closed-loop evaluation lets the planner's decisions affect the simulated scenario going forward. Increasingly that simulator is a NeRF or Gaussian-splatting reconstruction of a real log, or a learned world model, rather than something hand-authored.

Closed-loop is the only way to catch the failures that appear after the car has already made one small mistake.

04The math

Two equations carry almost all of the geometry on this page: the camera projection equation, and the epipolar constraint.

$$\lambda \begin{bmatrix}u\\v\\1\end{bmatrix} = K \big[R \mid t\big] \begin{bmatrix}X\\Y\\Z\\1\end{bmatrix}$$
  • X, Y, Z a point's coordinates in the world (or vehicle) frame.
  • R, t the camera's extrinsics: a rotation and translation converting that world-frame point into the camera's own coordinate frame.
  • K the camera's intrinsic matrix (focal length and principal point), projecting the camera-frame 3D point down onto the 2D image plane.
  • u, v the resulting pixel coordinates, up to the scale factor $\lambda$. That scale factor is the point's depth, and the projection alone throws it away. Hence depth estimation being a separate, hard problem.

The epipolar constraint, for two calibrated cameras viewing the same point:

$$x'^{\top} F x = 0$$
  • x, x' the same real-world point's pixel coordinates in the first and second camera respectively.
  • F the fundamental matrix, encoding the two cameras' relative geometry: their relative $R, t$ plus intrinsics.
  • Given $x$ and $F$, this equation defines a single line in the second image, the epipolar line, that $x'$ is guaranteed to lie on. This is why stereo matching needs only a 1D search along a line and not a full 2D image search.

05Why BEV/occupancy, and why modular over end-to-end

Fusing raw per-camera detections after the fact is the obvious first design. It is also a worse one. Each camera's detector has already thrown information away by the time it commits to a final box. So if two cameras jointly saw a partially-occluded object, but neither alone saw enough to be confident, late fusion never gets to combine their partial evidence. It was discarded before fusion ran.

Lifting features into a shared BEV grid before detecting anything keeps that partial evidence alive long enough to combine. This is why BEVFormer-style architectures displaced "detect per camera, then merge boxes" as the default.

Occupancy networks push the same logic further. A bounding-box vocabulary is a closed set of categories decided at training time, and the real world reliably produces obstacles outside that set. Predicting occupancy directly, instead of committing to "this is definitely a car, definitely this size," degrades gracefully on exactly the long-tail shapes a fixed category list cannot name.

The modular-versus-end-to-end question does not have a settled answer. Treating it as settled is a tell that someone has not thought about the actual tradeoff.

A modular pipeline keeps perception, prediction and planning as separate stages with separate losses. Every stage is independently testable. If the car braked unnecessarily, you can inspect whether perception mis-detected something, prediction mis-forecast a trajectory, or planning made a bad call given otherwise-correct inputs. Each stage validates against its own held-out data.

An end-to-end network optimized on one final task loss can in principle find a jointly better solution across all three at once: a slightly different detection behavior that happens to produce much better downstream plans. But that joint optimum arrives with no clean way to say why the network did what it did. Which is precisely the property regulators, safety cases and post-incident investigations need.

The tension that decides it

Global optimality versus stage-by-stage interpretability.

End-to-end can reach a better solution. Modular can tell you why it reached the one it did. For a safety-critical system, the second property is not a nice-to-have you trade away for accuracy.

That single tension is the actual reason modular pipelines remain the deployed standard, while end-to-end stays a very active, not-yet-dominant research direction.

06Complexity, failure modes, and what breaks

CameraLiDARRadar
Depth accuracyPoor alone; inferred, not measuredExcellent, directly measuredModerate, coarse resolution
VelocityInferred across framesInferred across framesMeasured directly (Doppler)
CostCheapestMost expensiveCheap
WeaknessBad in low light/glare; no direct depthSparse at range; expensive; weather-sensitiveCoarse spatial resolution; hard to classify objects

What shows up in production

  • Sensor miscalibration. Extrinsics drift out of alignment over time, from a bump, a temperature change, a minor collision. Every downstream fusion step then silently trusts a wrong camera-to-vehicle transform, producing subtly wrong 3D positions that look entirely plausible.
  • Long-tail object shapes that a bounding-box detector's training categories never covered. Occupancy networks are built to catch these, and pure bounding-box pipelines fail on them quietly instead of loudly.
  • Radar/camera disagreement in edge cases. A large stationary metal object can produce a weak or ambiguous radar return. If the fusion logic over-trusts radar's usually-reliable velocity signal, a real stationary hazard can be under-weighted.
  • Compounding error across stages. Prediction errors feed straight into planning as if they were ground truth. An overconfident, wrong trajectory forecast for another vehicle produces a planning decision that is optimal for the wrong predicted future.
Common misconception

"More sensors is strictly more robust." Not automatically. Fusion has to correctly handle sensors that disagree.

A fusion system with a bad prior, over-trusting one modality or working from a subtly miscalibrated extrinsic transform, can produce a worse combined estimate than the single best sensor would have given alone. Sensor diversity only helps if the fusion logic knows how much to trust each sensor under current conditions. That is a real design problem and does not come free with more data streams.

2026 status

BEV transformers and occupancy networks are the deployed standard for perception on modern multi-camera (and camera+LiDAR) stacks. Modular pipelines, meaning separate perception, prediction and planning stages, remain the deployed default for safety-critical systems. End-to-end driving is an active, fast-moving research area with promising benchmark results. It has not broadly displaced modular pipelines in safety-certified production systems, and claims that it has should be treated skeptically absent a specific, verifiable source.

07Build this

The projection equation in section 04 throws depth away in a single line. Recovering it from two photographs, with your own code, is what makes the rest of this page feel mechanical instead of magical.

Project Recover depth from a stereo pair, then break the calibration ~3 hours · NumPy

Take one rectified stereo pair from a public driving dataset and compute a disparity map by hand. Turn it into depth, then drop the result into the smallest possible bird's-eye-view grid. Two real photographs, about a hundred lines. Everything BEVFormer adds on top of this is learned. This part is not, and it is worth seeing that geometry alone already puts objects roughly where they belong.

  1. Download one rectified stereo pair and its calibration from a public benchmark. KITTI's stereo set is the usual choice. Rectified means corresponding pixels already sit on the same image row. Note the focal length $f$ in pixels and the baseline $b$ in metres.
  2. Write block matching yourself. For each pixel in the left image, slide a small window along the same row of the right image. Score each shift by sum of absolute differences and keep the best. That 1D search along a row is the epipolar constraint from section 04, cashed in. No cv2.StereoBM. Downscale the pair if the loop crawls.
  3. Convert with $\text{depth} = fb/d$ and show the disparity map beside the left image. Pick a few pixels by hand: a car just ahead, a car further down the road, a building at the end of the street. Compare their depths against each other.
  4. Back-project every pixel into vehicle-frame 3D by inverting the projection equation from section 04. Bin the points into a top-down grid of, say, 20 cm cells and plot the counts. That grid is a BEV with no learning in it at all.
  5. Now break the calibration. Change the baseline by a few percent and rerun everything, with the images untouched.
  6. Restore the baseline. Instead rotate the assumed camera extrinsic by a fraction of a degree, rerun, and look at the BEV again.
You'll know it worked when the disparity map is bright on near surfaces and dark on far ones, with edges where object edges are. The BEV grid should put a blob near the bottom for the car ahead, and a band at the far edge for the buildings. Nothing you wrote knows what a car is. That layout came out of geometry.
What the two breakages teach. A wrong baseline scales every depth by a constant, so the disparity map looks exactly as convincing as before and the BEV is uniformly wrong. Nothing in either picture flags it. A small extrinsic rotation is nastier. It tilts the whole point cloud, so the ground plane slopes and far points slide sideways much further than near ones. That is the miscalibration failure mode from section 06, reproduced in a few lines. Calibration errors do not look like errors. They look like confident, plausible, wrong geometry, which is why a real stack spends so much effort on staying calibrated.

Where this runs in production

A representative modern stack runs five stages:

  • Sense. Six to eight cameras around the vehicle, sometimes paired with LiDAR.
  • Fuse. A BEV transformer lifts every camera's features into one shared top-down grid, fusing across cameras and across the last few timesteps in the same step.
  • Detect and track. Running on the fused BEV grid, producing tracked objects with position, velocity and a class.
  • Predict. A trajectory predictor takes those tracks plus map context and forecasts several plausible short-horizon futures per agent.
  • Plan and execute. A planner searches candidate ego trajectories against all those predictions plus map and traffic-rule constraints. The winning trajectory goes to a classical trajectory-tracking controller (21).

Public benchmarks like nuScenes and the Waymo Open Dataset exist so different groups' perception and prediction stacks can be compared against the same real, sensor-synchronized driving logs, rather than each group's own private data.

08Interview questions

BeginnerWhy do self-driving cars need multiple sensor types instead of just one really good camera?

Because each sensor fails differently. Cameras are cheap and semantically rich but only give a 2D projection — depth has to be inferred, not measured, and they degrade badly in low light or glare. LiDAR measures 3D geometry directly and precisely, but it's expensive, sparse at range, and tells you nothing about color or texture. Radar measures velocity directly via Doppler shift — something neither camera nor LiDAR does without inferring it across multiple frames — but at coarse spatial resolution. No single sensor covers all three needs (rich semantics, precise geometry, direct velocity), so fusion isn't a nice-to-have, it's structurally required.

BeginnerWhat's the difference between a camera's intrinsics and extrinsics?

Intrinsics describe properties fixed by the camera itself — focal length and principal point — and map a 3D point already in the camera's own coordinate frame to a 2D pixel. Extrinsics describe where the camera is physically mounted and how it's oriented relative to some external reference frame — a rotation and translation converting a point from that external frame into the camera's frame, before intrinsics take over. A multi-camera rig needs every camera's extrinsics relative to one shared frame so their detections can be placed consistently relative to each other.

IntermediateHow does BEVFormer turn several camera images into one shared representation?

It defines a fixed grid of learned query vectors, one per cell of a top-down bird's-eye-view grid. Each query uses spatial cross-attention to pull in features from whichever camera views actually observe that patch of ground, using known camera geometry to know where to look. It also uses temporal self-attention against the previous timestep's BEV grid, which helps estimate velocity and recover briefly-occluded objects, since the model isn't reconstructing the scene from scratch every single frame. The result is one shared top-down feature grid that detection, tracking, and everything downstream can consume, instead of per-camera detections that have to be merged after the fact.

IntermediateWhy would you use an occupancy network instead of a standard 3D object detector?

A bounding-box detector assumes a fixed, closed set of object categories decided at training time, and real driving reliably produces obstacles that don't fit any of them — a fallen ladder, spilled cargo, debris. An occupancy network predicts, per grid cell, whether space is occupied and roughly what kind of thing occupies it, which degrades gracefully on unusual shapes instead of failing to detect them at all, because it never had to commit to a specific named category to register that the space is occupied.

DeepWhy hasn't the industry converged on end-to-end driving despite its promise of joint optimization?

End-to-end driving optimizes one network against a single end task loss, which can in principle discover jointly better solutions a modular pipeline's separately-trained stages can't see — a slightly different detection behavior that happens to produce a much better downstream plan, for instance. The cost is interpretability: when a modular pipeline makes a bad decision, you can inspect which stage — perception, prediction, or planning — was responsible, and validate each stage independently against held-out data for that specific sub-problem. An end-to-end network's decision doesn't decompose that way. For a safety-critical system that needs a debuggable, auditable answer to "why did the car do that," that loss of interpretability is a real cost, which is why modular pipelines remain the deployed default even though end-to-end methods post competitive results on research benchmarks.

DeepWhy is stereo depth estimation less reliable at long range than short range?

Depth is inversely proportional to disparity for a fixed camera baseline and focal length. At long range, disparity is already small — a handful of pixels — so a fixed amount of pixel-matching noise (say, one pixel of error) is a much larger relative error on a small disparity value than the same one-pixel error would be on the large disparity of a nearby object. That's a direct geometric consequence of the depth-disparity relationship, not a fixable calibration issue — it's why LiDAR, which measures range directly rather than through disparity, tends to be trusted more than stereo camera depth at longer distances.

09Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 20 3D, Spatial AI & Autonomous Driving

Free notes

Highlights on this page