Skip to main content
Open In Colab This section was written by Ruimeng Yang (pull request #58), with help from an AI coding agent on the code. It reproduces the ideas of Chapter 39 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. This section follows the book’s chapter on camera modeling and calibration. Each part pairs the book’s explanation, with a link to its figure, with a small experiment in PyTorch that makes the geometry concrete: projection at several depths, the effect of the intrinsic parameters, backprojecting a pixel to a ray, camera pitch and the horizon, and a DLT calibration on synthetic data. It ends with a real calibration attempt on a measured meeting room. The pinhole model, camera calibration, and Zhang’s method pages derive the same model and a practical calibration method in more depth.

Introduction

Chapter 39 shifts from the simple pinhole stories of earlier chapters to a more useful camera model: one that can describe where the camera sits in the world, how it is oriented, and how 3D points become pixels. Figure 39.1 in the book contrasts a camera-centered frame with a world-centered frame. It matters here because the whole chapter is about moving between those frames cleanly. Figure 39.2 in the book shows the image whose geometry the chapter later models. It matters because calibration always connects a real picture back to a 3D scene. The chapter’s core idea is that calibration is about estimating the transformation from world coordinates to image coordinates. Intrinsic parameters describe the camera itself. Extrinsic parameters describe the camera’s position and pose relative to the scene.

3D camera projections in homogeneous coordinates

Perspective projection contains a division by depth, which makes the equations awkward in ordinary Euclidean coordinates. Homogeneous coordinates express the same mapping as a matrix multiplication followed by a normalization step. Figure 39.3 in the book shows how a 3D point projects through a pinhole onto the image plane. It matters because homogeneous coordinates re-express this geometry in matrix form. The point to keep in mind is that a 3D location P = [X, Y, Z]^T projects to image coordinates roughly proportional to X/Z and Y/Z. Nearby points look larger because dividing by a smaller depth amplifies the image coordinates.

Parallel projection

Parallel projection removes that depth-dependent division. It is a simpler model that is sometimes useful for distant scenes or for deriving intuition, but it does not capture the foreshortening that real pinhole cameras produce.
Output from cell 8 Supplemental visualization. The figure projects the same lateral 3D point at three depths using both perspective and parallel projection. The perspective panel shrinks image coordinates with increasing depth, while the parallel panel keeps them fixed. That is the key geometric difference between perspective and parallel projection.

Camera-intrinsic parameters

Intrinsic parameters describe how the camera turns rays into pixels. In practice, this includes focal scaling, the principal point, and sometimes unequal pixel scales, skew, or lens distortion.

From meters to pixels

Figure 39.4 in the book shows how focal length, sensor width, and pixel sampling connect physical geometry to image coordinates. It matters because intrinsic calibration lives in that conversion. The intrinsic matrix converts camera-plane coordinates into pixel coordinates. The chapter uses a and b for horizontal and vertical focal scaling, and (c_x, c_y) for the principal point. Figure 39.5 in the book compares common image-coordinate conventions. It matters because sign choices and origin placement change how you write the intrinsic matrix. In code, it is worth being explicit about conventions: image origins, axis directions, and sign choices all affect the exact matrix form even when the geometry is the same.

From pixels to rays

Figure 39.6 in the book shows that a single pixel defines a whole 3D ray, not a unique 3D point. It matters because depth is what turns that ray back into a specific scene location. Backprojection reverses the forward camera map only up to a ray. If a pixel is known and the depth is unknown, there are infinitely many 3D points consistent with that pixel.

A simple, although unreliable, calibration method

Figure 39.7 in the book groups the physical setup and the captured chessboard image. It matters because the chapter first introduces calibration as a measurement problem before presenting more general estimation methods. This sanity-check method uses a known target, a measured distance, and the apparent size of the target in the image to estimate focal scaling. It is useful for intuition, but it ignores many practical effects such as distortion and imperfect alignment.

Other camera parameters

Real cameras may also need skew terms, separate horizontal and vertical pixel scales, and especially distortion correction. In practice, radial distortion is often the first extra effect that visibly breaks the simplest pinhole model.
Output from cell 10 Computational reconstruction related to Figure 39.4 of the book. The figure projects the same nine 3D points while changing either focal scaling or the principal point. Larger focal length magnifies the image, while changing (c_x, c_y) translates all projected points together. Those are two of the central effects encoded by the intrinsic matrix.
Output from cell 12 Computational reconstruction related to Figure 39.6 of the book. The code chooses one pixel direction and sampled three depths along its backprojected ray. The computed points all correspond to the same image location, which is exactly why a pixel alone is insufficient to recover a unique 3D point.

Camera-extrinsic parameters

Extrinsic parameters answer a different question from intrinsics: where is the camera in the world, and how is it rotated relative to the world frame? Figure 39.8 in the book shows the extrinsic relationship between the world frame and the camera frame. It matters because R and T explain where the camera is and how it is oriented. The chapter writes this mapping as a rotation followed by a translation in homogeneous coordinates. Once points are expressed in the camera frame, the intrinsic matrix can project them into the image.

Full camera model

Figure 39.9 in the book summarizes the full pipeline from world coordinates to camera coordinates to pixels. It matters because the projection matrix combines all of those steps. The full projection matrix composes intrinsics with extrinsics. In compact notation, the camera matrix is often written as M = K [R | -RT].

A few concrete examples

Figure 39.10 in the book collects the concrete camera-pose cases discussed in the chapter. It matters because the same matrix model can represent all of them. The chapter then walks through level cameras, tilted cameras, and more structured ground scenes to show how pose affects the final image equations. Figure 39.11 in the book shows a practical horizon cue in a level camera. It matters because camera orientation leaves visible traces in ordinary photographs. Figure 39.12 in the book sketches why equal-height points can line up in the image despite being at different depths. It matters because the horizon is a geometric consequence of camera pose. Figure 39.13 in the book groups two photos taken with different camera tilt angles. It matters because changing the camera pitch shifts the horizon line in a predictable way. A helpful way to read these examples is to ask which quantities stay fixed when the camera tilts or translates. Horizon-line motion is especially useful because it gives a visible signature of camera pitch.
Output from cell 14 Supplemental visualization. The figure projects a small set of equal-height scene points with a level camera and with a 15 degree downward pitch. The dashed line marks the numerically verified horizon height y_h = -f tan(theta) under the convention used here, P_C = R(P_W - T), and the displayed image coordinates use the standard image convention where y increases downward.

Camera calibration

Calibration estimates the mapping from known 3D scene points to observed 2D image points. The chapter first introduces a linear estimate of the projection matrix and then discusses how to recover more interpretable intrinsic and extrinsic parameters.

Direct linear transform

DLT solves for a projection matrix by stacking linear equations from several 3D-to-2D correspondences. The result is determined only up to an overall scale, which is fine for projective geometry.

Recovering intrinsic and extrinsic camera parameters

Once a projection matrix has been estimated, it still has to be factored into camera intrinsics and pose. Conceptually, this means separating the left 3x3 calibration part from the world-to-camera pose terms.

Multiplane calibration method

Multiplane calibration uses several views of a known planar target to stabilize the parameter estimates. This is closer to what practical camera-calibration toolkits do.

Nonlinear optimization by minimizing reprojection error

Figure 39.14 in the book shows reprojection error directly on the image plane. It matters because nonlinear refinement usually optimizes camera parameters by minimizing this quantity. Linear estimates are often only the starting point. A better model is usually obtained by refining parameters to minimize the distance between observed image points and predicted image points.

A toy example

Figure 39.15 in the book shows the real office scene used in the toy calibration example. It matters because the later 3D annotations are grounded in these measurements. Figure 39.16 in the book pairs image observations with measured 3D coordinates. It matters because calibration needs matched 3D scene points and 2D image points. Figure 39.17 in the book visualizes the inferred camera location from several viewpoints. It matters because a good calibration should produce a physically plausible camera pose. The office example shows the whole calibration story end to end: measure some 3D positions, annotate the corresponding image points, estimate the camera, and then inspect whether the result is plausible.
Output from cell 16
Computational reconstruction related to Figure 39.14 of the book. The code generates synthetic 3D points, projects them with a known camera, adds small image noise, builds the DLT linear system, and estimates a projection matrix up to scale. The left panel compares observed and predicted image points, the center panel shows the stacked DLT system, and the right panel summarizes reprojection error.

A real calibration attempt: a measured meeting room

This experiment, by the section’s author, mirrors the book’s office example (Figures 39.15 to 39.17) with a different scene: a photograph of a meeting room, scene measurements taken by hand in centimeters, hand-picked 3D-to-2D correspondences, a normalized DLT camera estimate, reprojection checks, and views of the measured scene. The calibration fails, and that is the lesson. The fitted camera reprojects the eight retained points with an RMS error of under 8 pixels, which looks good. But the recovered camera sits in an implausible place and looks away from the scene, and leaving out a single point moves the prediction by up to about 250 pixels. A low reprojection error on a handful of hand-measured points does not guarantee a physically correct camera. Real calibration needs many well-spread, non-coplanar correspondences, which is why practical methods use a calibration target seen from many views. The measurements use a local floor reference frame next to the planter, not a full architectural room model. The origin P1 is the left end of the green 61 cm floor annotation, immediately to the left of the planter. From that point:
  • +X follows the green 61 cm floor segment from P1 toward P3;
  • +Y follows the red 28 cm floor segment from P3 toward P4;
  • +Z points vertically upward.
So the floor is Z = 0. The blue television wall is modeled as the plane Y = -18 cm, because the cyan 18 cm floor offset places the wall reference behind P3. The world axes are mutually perpendicular, while their perspective projections in the photograph are generally not 90 degrees apart. The code for this experiment is long and mostly bookkeeping, so only its printed diagnostics and figures are shown.
measurementvalue_cmsource segmentstatus
Local X reference P1 -> P361.00Green 61 cm floor annotationdirect
Local Y reference P3 -> P428.00Red 28 cm floor annotationdirect
Local Z reference P3 -> P653.00Red 53 cm vertical planter annotationdirect
Wall offset from P3 to W018.00Cyan 18 cm floor offsetdirect
Wall height from W0 to P2300.00Yellow 300 cm wall annotationdirect
TV wall span beginning at X=61 on Y=-18193.00Green 193 cm annotationdirect
TV width along X148.00Green 148 cm annotation along the TV lower edgedirect
TV lower-left X coordinate83.50Derived as 61.0 + 22.5derived
TV lower-right X coordinate231.50Derived as 83.5 + 148.0derived
Table left X visualization coordinate131.00Derived as 61.0 + 70.0derived visualization assumption
Table back Y visualization coordinate80.00Derived as -18.0 + 98.0derived visualization assumption
Table width along X120.00Red 120 cm annotationdirect
Table length along Y240.00Cyan 240 cm annotationdirect
Table height along Z75.00Red 75 cm annotationdirect
idfeatureworld coordinateimage coordinatemeasurement endpointscoordinate derivationcalibration usestatusexclusion reason
1local floor origin(np.float64(0.0), np.float64(0.0), np.float64(0.0))(np.float64(18.0), np.float64(810.0))left endpoint of green 61 cm segmentchosen local referenceused by DLTmanually selected from physical endpoint-
3end of 61 cm X segment(np.float64(61.0), np.float64(0.0), np.float64(0.0))(np.float64(178.0), np.float64(790.0))green 61 cm segmentP1 + (61,0,0)used by DLTmanually selected from physical endpoint-
4end of 28 cm Y segment(np.float64(61.0), np.float64(28.0), np.float64(0.0))(np.float64(111.0), np.float64(858.0))red 28 cm segment from P3 to P4P3 + (0,28,0)used by DLTmanually selected from physical endpoint-
5occluded planter upper-left cornerunknownunavailablevertical above P3-side planter cornerunknown because the physical corner is occludedexcluded from DLTexcludedphysical corner is occluded or ambiguous
6top of 53 cm planter vertical(np.float64(61.0), np.float64(0.0), np.float64(53.0))(np.float64(166.0), np.float64(712.0))red 53 cm segment from P3 to P6P3 + (0,0,53)used by DLTmanually selected from physical endpoint-
20wall-floor reference at bottom of 300 cm segment(np.float64(61.0), np.float64(-18.0), np.float64(0.0))(np.float64(215.0), np.float64(721.0))cyan 18 cm offset and yellow 300 cm bottom endpointP3 + (0,-18,0)used by DLTmanually selected from physical endpoint-
2top of 300 cm wall reference(np.float64(61.0), np.float64(-18.0), np.float64(300.0))(np.float64(110.0), np.float64(181.0))yellow 300 cm segment top endpointW0 + (0,0,300)used by DLTmanually selected from physical endpoint; derived world coordinate-
7TV lower-left(np.float64(83.5), np.float64(-18.0), np.float64(110.0))(np.float64(224.0), np.float64(563.0))centered 148 cm width inside 193 cm span on wall planederived world coordinate on Y=-18 planeused by DLTmechanically valid derived world coordinate; pixel manually selected-
8TV lower-right(np.float64(231.5), np.float64(-18.0), np.float64(110.0))(np.float64(456.0), np.float64(551.0))centered 148 cm width inside 193 cm span on wall planederived world coordinate on Y=-18 planeused by DLTmechanically valid derived world coordinate; pixel manually selected-
9table back-left visualization point(np.float64(131.0), np.float64(80.0), np.float64(75.0))(np.float64(319.0), np.float64(633.0))61 cm reference, 70 cm X offset, 98 cm wall setback, 75 cm heightaxis-aligned visualization assumptionexcluded from DLTexcluded: orientation not measuredtable yaw relative to local axes not independently measured
10table back-right visualization point(np.float64(251.0), np.float64(80.0), np.float64(75.0))(np.float64(552.0), np.float64(617.0))120 cm table widthaxis-aligned visualization assumptionexcluded from DLTexcluded: orientation not measuredtable yaw relative to local axes not independently measured
11table front-left visualization point(np.float64(131.0), np.float64(320.0), np.float64(75.0))(np.float64(501.0), np.float64(908.0))240 cm table lengthaxis-aligned visualization assumptionexcluded from DLTexcluded: orientation not measuredtable yaw relative to local axes not independently measured
12table front-right visualization point(np.float64(251.0), np.float64(320.0), np.float64(75.0))unavailable120 cm width and 240 cm lengthaxis-aligned visualization assumptionexcluded from DLTexcluded: endpoint ambiguous and orientation not measuredoutside or not reliably visible
Output from cell 18 Output from cell 18 Output from cell 18
idobserved (u_i, v_i)predicted (u_i, v_i)fit error (px)leave-one-out error (px)
1(18.0, 810.0)(24.7, 819.3)11.4671.32
3(178.0, 790.0)(172.5, 780.0)11.3716.96
4(111.0, 858.0)(113.6, 860.5)3.61186.80
6(166.0, 712.0)(156.5, 709.9)9.7511.84
20(215.0, 721.0)(214.2, 723.1)2.299.85
2(110.0, 181.0)(108.0, 182.2)2.35246.95
7(224.0, 563.0)(231.7, 557.6)9.4028.76
8(456.0, 551.0)(457.0, 553.8)2.93126.87

Concluding remarks

A camera model is only useful when it links geometry, coordinates, and measurable image data. Intrinsics explain how rays become pixels, extrinsics explain where the camera is, and calibration ties both to real correspondences. The meeting-room experiment shows the last step is the fragile one: the correspondences must constrain the camera, not just fit it.