Skip to main content
Open In Colab This section was written by Ruimeng Yang (pull request #58), with help from an AI coding agent on the code. It reproduces the ideas of Chapter 40 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. This section follows the order of the book’s stereo chapter and turns it into executable experiments. Where the book uses a photograph or a diagram, a link points to it. Where geometry, matching, failure analysis, or evaluation can be computed, the code computes it on a synthetic stereo pair with known ground truth, so every result can be measured. Each topic follows the same steps: the concept, the relevant figure in the book, the math, an experiment, its visualization, a quantitative reading of the result, and the failure cases.

Introduction

Stereo begins as a perceptual phenomenon: the left and right eyes see slightly different image positions, and those displacements can be interpreted as depth. Figure 40.1 in the book: stereo anaglyph of the Titanic and the red/cyan viewing setup that turns left-right displacement into a 3D percept. The chapter quickly turns that perception into a two-part computational problem:
  1. geometry: where can a corresponding point lie?
  2. matching: which candidate point is the correct one?
The rest of this section follows the book’s order and adds executable experiments for the same ideas.

Stereo cues

How far away is a boat?

The boat example introduces two depth cues before stereo algorithms appear:
  • a single-view estimate using observer height h and angle alpha, with d = h / tan(alpha);
  • a two-view triangulation estimate using baseline t and angles alpha and beta, with d = t sin(alpha) sin(beta) / sin(alpha + beta).
Figure 40.2 in the book: two boat-distance constructions: one using a horizon reference and one using triangulation between two observation points.

Depth from image disparities

The random-dot stereogram isolates disparity from recognizable object identity. It shows that a depth percept can arise purely from left-right displacement. Figure 40.3 in the book: random-dot stereogram showing that disparity alone can create a depth percept.

Building a stereo pinhole camera

The chapter also grounds the perceptual story in image formation by showing a homemade anaglyph pinhole camera. Figure 40.4 in the book: an anaglyph pinhole camera built from two pinholes, color filters, and a projection plane. The next cell turns those ideas into five figures: geometric stereo cues, rectified stereo geometry, the inverse disparity-depth curve, far-depth sensitivity, and baseline sensitivity.
Output from cell 7 Output from cell 8
Output from cell 10
Output from cell 12
Output from cell 14
The figures above show, in order:
  • Boat triangulation and a random-dot stereogram make the book’s depth cue visible before any stereo algorithm appears.
  • Rectified stereo geometry with baseline, focal length, corresponding points, disparity, and a triangulated 3D point.
  • The inverse relationship between disparity and depth. Equal disparity steps do not correspond to equal depth steps.
  • A fixed disparity error creates much larger depth error at long range, which is why far geometry is fragile.
  • Larger baselines improve disparity signal but also create stronger view changes, overlap loss, and occlusion risk.
These figures sharpen three points from the book:
  • positive disparity means the corresponding point appears farther left in the right image;
  • depth is nonlinear in disparity;
  • far points are especially sensitive to small disparity errors, so geometry alone does not make stereo easy.

Model-based methods

Triangulation

The simple rectified geometry above is the special case used to derive the book’s core equations:
  • disparity: d = x_L - x_R
  • depth: Z = fB / d
Figure 40.5 in the book: rectified stereo geometry with focal length, baseline, image coordinates, and a triangulated 3D point.

Stereo matching

Once the geometry is fixed, the practical problem is correspondence. Which pixel in the right image matches a given pixel in the left? Figure 40.6 in the book: a real stereo office pair with corresponding features and their displacements highlighted. The next experiments use a synthetic rectified pair to compute a cost volume, run winner-takes-all disparity estimation, and measure actual error.
Output from cell 17 Output from cell 18
Output from cell 20
Output from cell 22
Output from cell 24
The figures above show, in order:
  • Single-pixel matching is brittle; patch aggregation stabilizes the cost surface, while brightness shifts still move the optimum.
  • Each disparity slice of the cost volume answers one question: how plausible is this disparity at every image location?
  • Winner-takes-all disparity estimation chooses d*(x,y) = argmin_d C(x,y,d) independently at each pixel.
  • Patch size is a real tradeoff: too small is noisy, too large blurs across discontinuities and occlusions.
  • The disparity search range must be large enough to include valid solutions but small enough to avoid wasted runtime and extra ambiguity.
Quantitatively, the experiments expose the actual matching objective: d*(x, y) = argmin_d C(x, y, d) where C(x, y, d) is the matching cost at pixel (x, y) for candidate disparity d. The sweeps make three failure modes visible:
  • too-small patches are unstable;
  • too-large patches bleed across depth discontinuities;
  • incorrect disparity ranges either truncate the solution or waste compute.

Finding image features

Feature-based stereo is the book’s answer to the fragility of raw intensity matching. Figure 40.8 in the book: feature detections on the office pair and the interpolated depth result. A good feature is localizable under small translations. The chapter expresses that with the Harris patch energy E(Delta x, Delta y) = sum_(x,y in P) (l(x,y) - l(x + Delta x, y + Delta y))^2 and then motivates richer local descriptors.

Local image descriptors

Figure 40.9 in the book: orientation-based local descriptors and the spatial pooling idea behind SIFT-style matching. Oriented local structure is often more stable than raw intensity values, especially under small view changes.

Interpolation between feature matches

Even after sparse feature matching, a system still needs interpolation or regularization to obtain dense depth. The next figures analyze the failure cases that make that interpolation necessary.
Output from cell 27 Output from cell 28
Output from cell 30
Output from cell 32
The figures above show, in order:
  • In a low-texture region the matching-cost surface becomes flat, so many disparities look almost equally plausible.
  • Repeated patterns create several near-identical alignments, which makes the cost surface multimodal.
  • Left-right consistency reveals pixels whose correspondences are unstable or absent because of occlusion.
  • A local parabola fit around the best integer disparity can recover a subpixel estimate when the cost curve is well behaved.
The failure modes, side by side:
  • textureless regions fail because the cost surface is flat;
  • repetitive textures fail because several disparities have similar cost;
  • occlusions fail because a point may not exist in both views at all;
  • brightness changes shift the entire cost curve, even when geometry is correct;
  • too-small patches are noisy;
  • too-large patches cross depth boundaries and mix objects;
  • far-depth sensitivity amplifies small disparity mistakes into large depth errors.
Practical mitigations exist, but each one leaves a tradeoff behind. Smoother costs reduce noise but blur discontinuities, larger baselines improve precision but increase occlusion, and subpixel fits help only when the local minimum is already reliable.

Constraints for arbitrary cameras

Once you leave the rectified special case, corresponding points are no longer found by searching the same row in the second image. Figure 40.10 in the book: candidate-match ambiguity for a point viewed under arbitrary camera geometry. Figure 40.11 in the book: the viewing ray from camera 1 projects to an epipolar line in camera 2. Figure 40.12 in the book: epipolar plane, epipolar lines, and epipoles for a stereo pair. The epipolar constraint is the algebraic form of that geometry: x'^T F x = 0 Valid correspondences should make the residual close to zero.

The essential and fundamental matrices

The essential matrix uses calibrated camera coordinates, while the fundamental matrix absorbs the intrinsic calibration and operates in image coordinates.

Epipolar lines: the game

Figure 40.13 in the book: epipolar-line intuition game matching camera arrangements to line families.

Image rectification

Rectification is a practical warp that makes epipolar lines horizontal again, reducing a 2D search problem to a 1D search problem.
Output from cell 35 Output from cell 36
The figures above show, in order:
  • The point-to-line relation becomes algebraic through the epipolar constraint x'^T F x = 0.
  • Conceptual rectification illustration: it shows the geometric goal of horizontal scanlines, not the output of a full calibrated rectification pipeline.
These figures reconnect the later epipolar geometry back to the earlier rectified stereo experiments. Rectification is not a different problem; it is a practical reparameterization of the same correspondence constraint.

Learning-based methods

The book does not present a full deep-learning survey. Instead, it emphasizes why many learned systems still predict disparity rather than depth directly, and why they often use a rectified cost-volume pipeline. Figure 40.14 in the book: two-stage CNN stereo pipeline: feature extraction, cost volume, cost aggregation, and disparity estimate. The ideas to keep from it:
  • disparity is easier to regularize than depth under a fixed stereo rig;
  • cost volumes remain central even in learned stereo;
  • the network predicts disparity, but geometry still interprets the result.

Evaluation

The evaluation is executable: it measures classical block matching on the synthetic stereo pair.
Output from cell 39
Output from cell 41
The figures above show, in order:
  • Evaluation should expose both disparity-space and depth-space error, because small disparity errors can turn into large depth errors far from the cameras.
  • Runtime and accuracy should be read together. The timings are machine-dependent and are intended only for relative comparison within this run.
The code reports these metrics:
  • disparity MAE: mean absolute disparity error over valid visible pixels;
  • bad-pixel rate: fraction of valid pixels with absolute disparity error greater than 1 pixel;
  • valid-pixel ratio: fraction of all pixels that remain usable after support and visibility checks;
  • depth RMSE: root-mean-square error after converting disparity to depth;
  • left-right consistency rate: fraction of valid pixels that agree with a reverse-direction disparity check;
  • runtime: wall-clock runtime for the selected matching configuration.

Concluding remarks

The book ends by staying realistic about stereo:
  • correspondence is hard even in the rectified case;
  • far geometry is fragile because depth is nonlinear in disparity;
  • occlusions and repeated structure create genuine ambiguity;
  • practical systems mix geometry, matching, regularization, and evaluation rather than relying on a single elegant formula.
The experiments in this section reflect that balance: they run the book’s ideas as code and measure the consequences directly.