Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #75), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 42 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead.
A single photograph looks like it has lost all depth, yet with a few geometric facts you can measure a room from it: the heights of objects, the camera’s focal length, and the 3D position of points on the floor. This section builds every one of those measurements and checks each against ground truth.

A synthetic office, end to end

The book works on a photograph of an office (Figure 42.1). Here you build a synthetic stand-in: a small office scene written as a list of 3D line segments. Each segment is one edge of the room, the bookshelf, the desk, or the bottle, and world coordinates are in centimeters. The camera looks into the room from a known position with known intrinsics. All of this is ground truth that the rest of the section recovers using only the projected 2D image.
Output from cell 5

Linear perspective

A 3D line P(t)=P0+t D\mathbf{P}(t) = \mathbf{P}_0 + t\,\mathbf{D} projects to a 2D line that, as t→∞t \to \infty, ends at the vanishing point v=f (DX/DZ,  DY/DZ)\mathbf{v} = f\,\big(D_X / D_Z, \; D_Y / D_Z\big) (equation 42.3). Equation 42.5 generalizes this through the full camera matrix: v=KRD.\mathbf{v} = \mathbf{K}\mathbf{R}\mathbf{D}. Two consequences come up again and again:
  1. The vanishing point depends only on the line’s direction, not on where it starts. So all lines that are parallel in 3D meet at the same image point.
  2. All vanishing points of lines lying in a single plane lie on the plane’s horizon line.
Output from cell 6 Output from cell 7

Detecting vanishing points

Algorithm 1. Given a set of 2D line segments believed to be parallel in 3D:
  1. For each pair of segments, compute their intersection in the image plane (the cross product of their homogeneous line vectors).
  2. Run RANSAC over those candidate intersections: a vanishing point is a location with many votes.
  3. Repeat for each direction. The office has three orthogonal world directions, so it has three vanishing points.
The code below runs this on the projected office scene. Because the ground-truth vanishing points are known (from KRD\mathbf{K}\mathbf{R}\mathbf{D}), you can check the recovered ones.
Output from cell 10 The book draws the same construction on its office photograph in Figure 42.10.

Measuring heights with the cross-ratio

For any four collinear points the cross-ratio is a projective invariant: it survives perspective projection unchanged (equation 42.6): CR(P1,P2,P3,P4)=∣P3−P1∣ ∣P4−P2∣∣P3−P2∣ ∣P4−P1∣.\mathrm{CR}(\mathbf{P}_1, \mathbf{P}_2, \mathbf{P}_3, \mathbf{P}_4) = \frac{|\mathbf{P}_3 - \mathbf{P}_1|\,|\mathbf{P}_4 - \mathbf{P}_2|}{|\mathbf{P}_3 - \mathbf{P}_2|\,|\mathbf{P}_4 - \mathbf{P}_1|}. This one invariant powers single-view metrology. The next cells demonstrate it, then use it (Algorithm 2 in the book) to measure the desk’s height from the projected image alone, given only that the bookshelf is 197 cm tall. Output from cell 11

Algorithm 2: measuring the desk’s height

From the chapter, given image points gb,tb\mathbf{g}_b, \mathbf{t}_b (bookshelf bottom/top), gd,td\mathbf{g}_d, \mathbf{t}_d (desk bottom/top), horizon line h\mathbf{h}, and vertical vanishing point v3\mathbf{v}_3: l1=gd×gb,a=l1×h,l2=a×td,l3=gb×tb,b=l2×l3.\mathbf{l}_1 = \mathbf{g}_d \times \mathbf{g}_b, \quad \mathbf{a} = \mathbf{l}_1 \times \mathbf{h}, \quad \mathbf{l}_2 = \mathbf{a} \times \mathbf{t}_d, \quad \mathbf{l}_3 = \mathbf{g}_b \times \mathbf{t}_b, \quad \mathbf{b} = \mathbf{l}_2 \times \mathbf{l}_3. Then the desk height comes out of the cross-ratio identity: hbookshelfhdesk=∣b−v3∣ ∣tb−gb∣∣b−gb∣ ∣tb−v3∣.\frac{h_{\text{bookshelf}}}{h_{\text{desk}}} = \frac{|\mathbf{b}-\mathbf{v}_3|\,|\mathbf{t}_b - \mathbf{g}_b|}{|\mathbf{b}-\mathbf{g}_b|\,|\mathbf{t}_b - \mathbf{v}_3|}.
Output from cell 14 The book runs Algorithm 2 on its real office photograph, marking the four image points and the 197 cm bookshelf (Figures 42.13 and 42.14).

Algorithm 3: height propagation to supported objects

Once you know the desk’s height, the bottle resting on it becomes measurable too. You project the bottle’s height onto the bookshelf with the same construction, but take the desk top as the reference instead of the floor: hbottle=hbookshelf∣c−gb∣ ∣tb−v3∣∣c−v3∣ ∣tb−gb∣−hdesk.h_{\text{bottle}} = h_{\text{bookshelf}} \frac{|\mathbf{c}-\mathbf{g}_b|\,|\mathbf{t}_b - \mathbf{v}_3|}{|\mathbf{c}-\mathbf{v}_3|\,|\mathbf{t}_b - \mathbf{g}_b|} - h_{\text{desk}}. The true bottle height in the synthetic scene is 25 cm.
Output from cell 16 On its real photograph the book recovers a bottle height of about 27.6 cm against an actual 25.5 cm (Figure 42.16).

Algorithm 4: axis calibration via cross-ratio

Given a vanishing point v\mathbf{v} on a calibrated axis, the origin o\mathbf{o} in the image, and an arbitrary reference point r\mathbf{r} at world-distance α\alpha from the origin, the cross-ratio determines the image position of every other tick: CR(o,r,rk,v)=k\mathrm{CR}(\mathbf{o}, \mathbf{r}, \mathbf{r}_k, \mathbf{v}) = k where rk\mathbf{r}_k is the projected position of kαk\alpha. Solving for rk\mathbf{r}_k given o,r,v,k\mathbf{o}, \mathbf{r}, \mathbf{v}, k gives the recipe.
Output from cell 18

Algorithm 5: locate a 3D point

Suppose you know a pixel p\mathbf{p} lies on the ground plane, for example at the base of a chair. You can recover its 3D world coordinate by reading off the calibrated world-axis tick that the perpendicular from p\mathbf{p} to the axis would hit. Equivalently, the back-projected ray from the camera through p\mathbf{p} meets the known ground plane at exactly one 3D point, and that point is the recovered world position.
Output from cell 20 Output from cell 20

Camera calibration from three vanishing points

Algorithm 6. If the scene contains three mutually orthogonal directions with vanishing points v1,v2,v3\mathbf{v}_1, \mathbf{v}_2, \mathbf{v}_3, then orthogonality gives three linear constraints on W=K−⊤K−1\mathbf{W} = \mathbf{K}^{-\top}\mathbf{K}^{-1} (equation 42.9): vi⊤W vj=0for (i,j)∈{(1,2),(1,3),(2,3)}.\mathbf{v}_i^\top \mathbf{W}\, \mathbf{v}_j = 0 \quad \text{for } (i, j) \in \{(1,2), (1,3), (2,3)\}. Solve for the four free parameters of W\mathbf{W} by SVD, then Cholesky-factor W=L⊤L\mathbf{W} = \mathbf{L}^\top\mathbf{L} to recover K=L−1\mathbf{K} = \mathbf{L}^{-1}, normalized so K33=1\mathbf{K}_{33} = 1. On synthetic data you can check that the recovered K\mathbf{K} matches K_true.
The book applies this calibration to its office photograph and draws the calibrated world axes and a measured box over it (Figures 42.17, 42.18, and 42.19). Output from cell 22

Concluding remarks

A single view gives you three kinds of measurement:
  • Three vanishing points give the camera intrinsics, through the SVD of the orthogonality constraints on W=K−⊤K−1\mathbf{W} = \mathbf{K}^{-\top}\mathbf{K}^{-1} and a Cholesky factorization. On the synthetic office, K\mathbf{K} comes back exactly.
  • The cross-ratio along a vertical line gives the height of any object standing on the floor (Algorithm 2) or on a supporting plane (Algorithm 3). The desk comes back at 76 cm and the bottle at 25 cm, both matching the synthetic ground truth.
  • The horizon line and calibrated world axes give the 3D position of any point on a known plane (Figure 42.20). Without a plane to anchor it, a single view leaves depth ambiguous.
This section does not cover the book’s perceptual demonstrations, such as the Ames room and the leaning towers: they are illusions for the eye, not algorithms.