Introduction
Chapter 39 shifts from the simple pinhole stories of earlier chapters to a more useful camera model: one that can describe where the camera sits in the world, how it is oriented, and how 3D points become pixels. Figure 39.1 in the book contrasts a camera-centered frame with a world-centered frame. It matters here because the whole chapter is about moving between those frames cleanly. Figure 39.2 in the book shows the image whose geometry the chapter later models. It matters because calibration always connects a real picture back to a 3D scene. The chapter’s core idea is that calibration is about estimating the transformation from world coordinates to image coordinates. Intrinsic parameters describe the camera itself. Extrinsic parameters describe the camera’s position and pose relative to the scene.3D camera projections in homogeneous coordinates
Perspective projection contains a division by depth, which makes the equations awkward in ordinary Euclidean coordinates. Homogeneous coordinates express the same mapping as a matrix multiplication followed by a normalization step. Figure 39.3 in the book shows how a 3D point projects through a pinhole onto the image plane. It matters because homogeneous coordinates re-express this geometry in matrix form. The point to keep in mind is that a 3D locationP = [X, Y, Z]^T projects to image
coordinates roughly proportional to X/Z and Y/Z. Nearby points look larger because
dividing by a smaller depth amplifies the image coordinates.
Parallel projection
Parallel projection removes that depth-dependent division. It is a simpler model that is sometimes useful for distant scenes or for deriving intuition, but it does not capture the foreshortening that real pinhole cameras produce.
Camera-intrinsic parameters
Intrinsic parameters describe how the camera turns rays into pixels. In practice, this includes focal scaling, the principal point, and sometimes unequal pixel scales, skew, or lens distortion.From meters to pixels
Figure 39.4 in the book shows how focal length, sensor width, and pixel sampling connect physical geometry to image coordinates. It matters because intrinsic calibration lives in that conversion. The intrinsic matrix converts camera-plane coordinates into pixel coordinates. The chapter usesa and b for horizontal and vertical focal scaling, and
(c_x, c_y) for the principal point.
Figure 39.5 in the book compares common image-coordinate conventions. It matters because sign choices and origin placement change how you write the intrinsic matrix.
In code, it is worth being explicit about conventions: image origins, axis directions,
and sign choices all affect the exact matrix form even when the geometry is the same.
From pixels to rays
Figure 39.6 in the book shows that a single pixel defines a whole 3D ray, not a unique 3D point. It matters because depth is what turns that ray back into a specific scene location. Backprojection reverses the forward camera map only up to a ray. If a pixel is known and the depth is unknown, there are infinitely many 3D points consistent with that pixel.A simple, although unreliable, calibration method
Figure 39.7 in the book groups the physical setup and the captured chessboard image. It matters because the chapter first introduces calibration as a measurement problem before presenting more general estimation methods. This sanity-check method uses a known target, a measured distance, and the apparent size of the target in the image to estimate focal scaling. It is useful for intuition, but it ignores many practical effects such as distortion and imperfect alignment.Other camera parameters
Real cameras may also need skew terms, separate horizontal and vertical pixel scales, and especially distortion correction. In practice, radial distortion is often the first extra effect that visibly breaks the simplest pinhole model.
(c_x, c_y) translates all projected points together. Those are two of the central effects encoded by the intrinsic matrix.

Camera-extrinsic parameters
Extrinsic parameters answer a different question from intrinsics: where is the camera in the world, and how is it rotated relative to the world frame? Figure 39.8 in the book shows the extrinsic relationship between the world frame and the camera frame. It matters becauseR and T explain where the camera is and how it is oriented.
The chapter writes this mapping as a rotation followed by a translation in homogeneous
coordinates. Once points are expressed in the camera frame, the intrinsic matrix can
project them into the image.
Full camera model
Figure 39.9 in the book summarizes the full pipeline from world coordinates to camera coordinates to pixels. It matters because the projection matrix combines all of those steps. The full projection matrix composes intrinsics with extrinsics. In compact notation, the camera matrix is often written asM = K [R | -RT].
A few concrete examples
Figure 39.10 in the book collects the concrete camera-pose cases discussed in the chapter. It matters because the same matrix model can represent all of them. The chapter then walks through level cameras, tilted cameras, and more structured ground scenes to show how pose affects the final image equations. Figure 39.11 in the book shows a practical horizon cue in a level camera. It matters because camera orientation leaves visible traces in ordinary photographs. Figure 39.12 in the book sketches why equal-height points can line up in the image despite being at different depths. It matters because the horizon is a geometric consequence of camera pose. Figure 39.13 in the book groups two photos taken with different camera tilt angles. It matters because changing the camera pitch shifts the horizon line in a predictable way. A helpful way to read these examples is to ask which quantities stay fixed when the camera tilts or translates. Horizon-line motion is especially useful because it gives a visible signature of camera pitch.
y_h = -f tan(theta) under the convention used here, P_C = R(P_W - T), and the displayed image coordinates use the standard image convention where y increases downward.
Camera calibration
Calibration estimates the mapping from known 3D scene points to observed 2D image points. The chapter first introduces a linear estimate of the projection matrix and then discusses how to recover more interpretable intrinsic and extrinsic parameters.Direct linear transform
DLT solves for a projection matrix by stacking linear equations from several 3D-to-2D correspondences. The result is determined only up to an overall scale, which is fine for projective geometry.Recovering intrinsic and extrinsic camera parameters
Once a projection matrix has been estimated, it still has to be factored into camera intrinsics and pose. Conceptually, this means separating the left3x3 calibration part
from the world-to-camera pose terms.
Multiplane calibration method
Multiplane calibration uses several views of a known planar target to stabilize the parameter estimates. This is closer to what practical camera-calibration toolkits do.Nonlinear optimization by minimizing reprojection error
Figure 39.14 in the book shows reprojection error directly on the image plane. It matters because nonlinear refinement usually optimizes camera parameters by minimizing this quantity. Linear estimates are often only the starting point. A better model is usually obtained by refining parameters to minimize the distance between observed image points and predicted image points.A toy example
Figure 39.15 in the book shows the real office scene used in the toy calibration example. It matters because the later 3D annotations are grounded in these measurements. Figure 39.16 in the book pairs image observations with measured 3D coordinates. It matters because calibration needs matched 3D scene points and 2D image points. Figure 39.17 in the book visualizes the inferred camera location from several viewpoints. It matters because a good calibration should produce a physically plausible camera pose. The office example shows the whole calibration story end to end: measure some 3D positions, annotate the corresponding image points, estimate the camera, and then inspect whether the result is plausible.
A real calibration attempt: a measured meeting room
This experiment, by the section’s author, mirrors the book’s office example (Figures 39.15 to 39.17) with a different scene: a photograph of a meeting room, scene measurements taken by hand in centimeters, hand-picked 3D-to-2D correspondences, a normalized DLT camera estimate, reprojection checks, and views of the measured scene. The calibration fails, and that is the lesson. The fitted camera reprojects the eight retained points with an RMS error of under 8 pixels, which looks good. But the recovered camera sits in an implausible place and looks away from the scene, and leaving out a single point moves the prediction by up to about 250 pixels. A low reprojection error on a handful of hand-measured points does not guarantee a physically correct camera. Real calibration needs many well-spread, non-coplanar correspondences, which is why practical methods use a calibration target seen from many views. The measurements use a local floor reference frame next to the planter, not a full architectural room model. The originP1 is the left end of the green 61 cm floor annotation, immediately to the left of the planter. From that point:
+Xfollows the green61 cmfloor segment fromP1towardP3;+Yfollows the red28 cmfloor segment fromP3towardP4;+Zpoints vertically upward.
Z = 0. The blue television wall is modeled as the plane Y = -18 cm, because the cyan 18 cm floor offset places the wall reference behind P3. The world axes are mutually perpendicular, while their perspective projections in the photograph are generally not 90 degrees apart. The code for this experiment is long and mostly bookkeeping, so only its printed diagnostics and figures are shown.
| measurement | value_cm | source segment | status |
|---|---|---|---|
| Local X reference P1 -> P3 | 61.00 | Green 61 cm floor annotation | direct |
| Local Y reference P3 -> P4 | 28.00 | Red 28 cm floor annotation | direct |
| Local Z reference P3 -> P6 | 53.00 | Red 53 cm vertical planter annotation | direct |
| Wall offset from P3 to W0 | 18.00 | Cyan 18 cm floor offset | direct |
| Wall height from W0 to P2 | 300.00 | Yellow 300 cm wall annotation | direct |
| TV wall span beginning at X=61 on Y=-18 | 193.00 | Green 193 cm annotation | direct |
| TV width along X | 148.00 | Green 148 cm annotation along the TV lower edge | direct |
| TV lower-left X coordinate | 83.50 | Derived as 61.0 + 22.5 | derived |
| TV lower-right X coordinate | 231.50 | Derived as 83.5 + 148.0 | derived |
| Table left X visualization coordinate | 131.00 | Derived as 61.0 + 70.0 | derived visualization assumption |
| Table back Y visualization coordinate | 80.00 | Derived as -18.0 + 98.0 | derived visualization assumption |
| Table width along X | 120.00 | Red 120 cm annotation | direct |
| Table length along Y | 240.00 | Cyan 240 cm annotation | direct |
| Table height along Z | 75.00 | Red 75 cm annotation | direct |
| id | feature | world coordinate | image coordinate | measurement endpoints | coordinate derivation | calibration use | status | exclusion reason |
|---|---|---|---|---|---|---|---|---|
| 1 | local floor origin | (np.float64(0.0), np.float64(0.0), np.float64(0.0)) | (np.float64(18.0), np.float64(810.0)) | left endpoint of green 61 cm segment | chosen local reference | used by DLT | manually selected from physical endpoint | - |
| 3 | end of 61 cm X segment | (np.float64(61.0), np.float64(0.0), np.float64(0.0)) | (np.float64(178.0), np.float64(790.0)) | green 61 cm segment | P1 + (61,0,0) | used by DLT | manually selected from physical endpoint | - |
| 4 | end of 28 cm Y segment | (np.float64(61.0), np.float64(28.0), np.float64(0.0)) | (np.float64(111.0), np.float64(858.0)) | red 28 cm segment from P3 to P4 | P3 + (0,28,0) | used by DLT | manually selected from physical endpoint | - |
| 5 | occluded planter upper-left corner | unknown | unavailable | vertical above P3-side planter corner | unknown because the physical corner is occluded | excluded from DLT | excluded | physical corner is occluded or ambiguous |
| 6 | top of 53 cm planter vertical | (np.float64(61.0), np.float64(0.0), np.float64(53.0)) | (np.float64(166.0), np.float64(712.0)) | red 53 cm segment from P3 to P6 | P3 + (0,0,53) | used by DLT | manually selected from physical endpoint | - |
| 20 | wall-floor reference at bottom of 300 cm segment | (np.float64(61.0), np.float64(-18.0), np.float64(0.0)) | (np.float64(215.0), np.float64(721.0)) | cyan 18 cm offset and yellow 300 cm bottom endpoint | P3 + (0,-18,0) | used by DLT | manually selected from physical endpoint | - |
| 2 | top of 300 cm wall reference | (np.float64(61.0), np.float64(-18.0), np.float64(300.0)) | (np.float64(110.0), np.float64(181.0)) | yellow 300 cm segment top endpoint | W0 + (0,0,300) | used by DLT | manually selected from physical endpoint; derived world coordinate | - |
| 7 | TV lower-left | (np.float64(83.5), np.float64(-18.0), np.float64(110.0)) | (np.float64(224.0), np.float64(563.0)) | centered 148 cm width inside 193 cm span on wall plane | derived world coordinate on Y=-18 plane | used by DLT | mechanically valid derived world coordinate; pixel manually selected | - |
| 8 | TV lower-right | (np.float64(231.5), np.float64(-18.0), np.float64(110.0)) | (np.float64(456.0), np.float64(551.0)) | centered 148 cm width inside 193 cm span on wall plane | derived world coordinate on Y=-18 plane | used by DLT | mechanically valid derived world coordinate; pixel manually selected | - |
| 9 | table back-left visualization point | (np.float64(131.0), np.float64(80.0), np.float64(75.0)) | (np.float64(319.0), np.float64(633.0)) | 61 cm reference, 70 cm X offset, 98 cm wall setback, 75 cm height | axis-aligned visualization assumption | excluded from DLT | excluded: orientation not measured | table yaw relative to local axes not independently measured |
| 10 | table back-right visualization point | (np.float64(251.0), np.float64(80.0), np.float64(75.0)) | (np.float64(552.0), np.float64(617.0)) | 120 cm table width | axis-aligned visualization assumption | excluded from DLT | excluded: orientation not measured | table yaw relative to local axes not independently measured |
| 11 | table front-left visualization point | (np.float64(131.0), np.float64(320.0), np.float64(75.0)) | (np.float64(501.0), np.float64(908.0)) | 240 cm table length | axis-aligned visualization assumption | excluded from DLT | excluded: orientation not measured | table yaw relative to local axes not independently measured |
| 12 | table front-right visualization point | (np.float64(251.0), np.float64(320.0), np.float64(75.0)) | unavailable | 120 cm width and 240 cm length | axis-aligned visualization assumption | excluded from DLT | excluded: endpoint ambiguous and orientation not measured | outside or not reliably visible |



| id | observed (u_i, v_i) | predicted (u_i, v_i) | fit error (px) | leave-one-out error (px) |
|---|---|---|---|---|
| 1 | (18.0, 810.0) | (24.7, 819.3) | 11.46 | 71.32 |
| 3 | (178.0, 790.0) | (172.5, 780.0) | 11.37 | 16.96 |
| 4 | (111.0, 858.0) | (113.6, 860.5) | 3.61 | 186.80 |
| 6 | (166.0, 712.0) | (156.5, 709.9) | 9.75 | 11.84 |
| 20 | (215.0, 721.0) | (214.2, 723.1) | 2.29 | 9.85 |
| 2 | (110.0, 181.0) | (108.0, 182.2) | 2.35 | 246.95 |
| 7 | (224.0, 563.0) | (231.7, 557.6) | 9.40 | 28.76 |
| 8 | (456.0, 551.0) | (457.0, 553.8) | 2.93 | 126.87 |

