Skip to main content
Open In Colab This section was written by Kimberly Milner (pull request #44), with help from an AI coding agent on the code. It reproduces the ideas of Chapter 6 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. This section answers a question pinholes cannot: how do you let in more light without sacrificing sharpness? The answer starts with Snell’s law. The first part derives the lensmaker’s formula, which turns Snell’s law plus a small-angle approximation into one equation, 1a+1b=1f,\frac{1}{a} + \frac{1}{b} = \frac{1}{f}, relating object distance aa, image distance bb, and focal length ff. Each figure is a runnable construction: change a refractive index, a radius of curvature, or an object distance, and the diagram redraws from the new physics. The book’s diagrams are regenerated directly. Where the book shows a photograph, either a synthetic version computed from the same physics stands in, or a link points to the book’s figure.

Introduction

A pinhole camera works: light from a scene passes through a small opening and forms an image on a sensor on the other side. But pinholes force a trade-off. Shrink the aperture and each scene point maps to a tight spot: the image is sharp, but so little light gets through that the image is dim. Open the aperture and more light arrives, but each scene point now spreads across many sensor pixels: the image is bright but blurry. There’s no single pinhole size that’s both sharp and bright. A lens breaks the trade-off. It gathers the wide cone of light a large aperture admits and refocuses it back to a single point on the sensor: bright like the wide pinhole, sharp like the narrow one. The rest of this section is the geometry of how a lens does that. Figure 6.1: Brightness/sharpness trade-offs in image formation. Three setups, same scene at the top, same sensor at the bottom. (a) Small pinhole: sharp image, very dim: little light reaches the sensor. (b) Large pinhole: bright image, very blurry: each scene point spreads to a wide disk at the sensor. (c) Lens: bright AND sharp: the lens collects a wide cone of light from each scene point and refocuses it to one sensor point. The book’s version is Figure 6.1; it demonstrates the same tradeoff with photographs in Figure 6.2. Output from cell 7 The next section derives the geometry behind panel (c): how exactly does a lens refocus a wide cone of light from each scene point back to a single sensor point? The answer is Snell’s law applied twice (once at each glass surface), with a small-angle approximation that linearizes the algebra. That’s the lensmaker’s formula, which the next part builds.

The lensmaker’s formula

A lens refracts light at each of its two surfaces. The lensmaker’s formula compresses that two-refraction process into one equation, 1a+1b=1f,\frac{1}{a} + \frac{1}{b} = \frac{1}{f}, relating object distance aa, image distance bb, and focal length ff. It follows from applying Snell’s law twice with the small-angle approximation sin⁡θ≈θ\sin\theta \approx \theta: the paraxial regime where the algebra stays linear. For small angles in radians, sin⁡θ≈θ\sin\theta \approx \theta, and Snell’s law at a glass-air interface (n1=1n_1 = 1, n2=nn_2 = n) becomes the linear θ1=nθ2\theta_1 = n\theta_2. The book uses this paraxial form throughout the rest of the section. Figure 6.3(a): Snell’s law at a flat interface. A ray crosses from a medium with refractive index n1n_1 into a denser medium with n2>n1n_2 > n_1 and bends toward the normal. The angles θ1,θ2\theta_1, \theta_2 are measured from the normal, and n1sin⁡θ1=n2sin⁡θ2n_1 \sin\theta_1 = n_2 \sin\theta_2 relates them. The book’s panel (b), a photograph of a straw refracting in a glass of water, is the same physics in the physical world; that photo is not reproduced here. (Book Figure 6.3.) Output from cell 8 Figure 6.4(a): Thin-lens geometry. A point on the optical axis at distance aa from the lens emits rays in many directions. Those rays pass through the lens at various heights up to cc and converge to a single image point at distance bb on the other side. The angle the upper extreme ray makes with the optical axis is θ1\theta_1 on the object side and θ4\theta_4 on the image side. The lensmaker’s formula derived in this section is the statement that for a thin lens, this convergence happens at the same bb for every ray emitted from the same object point. (Book Figure 6.4(a).) Output from cell 9 The figure depicts an idealization: rays appear to bend at a single point (the center of the lens) rather than refracting twice (at each glass surface). This is the thin-lens approximation, which assumes the lens’s physical thickness is small enough to ignore. Panel (b), drawn next, distorts the geometry to expose the actual two-surface bending that the thin-lens approximation glosses over. Figure 6.4(b): The labeled thin-lens geometry. Panel (b) follows a single ray through the same thin lens as panel (a), with the geometry deliberately distorted, two surfaces pulled apart, angles enlarged, so every label stays legible. The ray refracts at the front surface, crosses the glass, refracts again at the back, and meets the axis at the image point, turning through the four angles θ1,θ2,θ3,θ4\theta_1, \theta_2, \theta_3, \theta_4. At each surface, the local tilt θS\theta_S is set by the height cc at which the ray crosses (brackets C1≈C2≈CC_1 \approx C_2 \approx C because the lens is thin, d≈0d \approx 0). The object distance aa, image distance bb, and negligible thickness d≈0d \approx 0 are marked below the axis. The paraxial Snell’s law at each surface gives the two relations the derivation sums: n θ2=θ1+θS,n θ3=θ4+θS.n\,\theta_2 = \theta_1 + \theta_S, \qquad n\,\theta_3 = \theta_4 + \theta_S. (Book Figure 6.4(b) and Table 6.1.) Output from cell 10 From the diagram to the lensmaker’s formula. The book’s Table 6.1 sums the four small-angle Snell relations along the ray’s path. Substituting the axis angles θ1≈c/a\theta_1 \approx c/a and θ4≈c/b\theta_4 \approx c/b, and the surface tilts θS1≈c/R1\theta_{S_1} \approx c/R_1 and θS2≈c/R2\theta_{S_2} \approx c/R_2, gives 1a+1b=(n−1)(1R1+1R2).\frac{1}{a} + \frac{1}{b} = (n-1)\left(\frac{1}{R_1} + \frac{1}{R_2}\right). The height cc cancels on both sides: every ray from the same object point lands at the same image distance bb, regardless of where it crosses the lens. Defining 1f=(n−1)(1R1+1R2)\frac{1}{f} = (n-1)\left(\frac{1}{R_1} + \frac{1}{R_2}\right) gives the lensmaker’s formula in the target form: 1a+1b=1f.\frac{1}{a} + \frac{1}{b} = \frac{1}{f}. (Book Figure 6.4(b) and Table 6.1.)

From flat interface to curved surface

The angles in the lensmaker derivation are measured from each lens surface’s normal, but the lens’s surfaces aren’t flat: they’re spherical. Two things change with curvature. The normal at the point where a ray strikes the surface is no longer parallel to the optical axis, and the tilt of that normal depends on how far above the axis the ray hits. The next figure pins down that dependence: how the surface’s tilt angle θS\theta_S relates to the radius of curvature RR and the hit height cc. Figure 6.5: Relation between RR and θS\theta_S. A spherical surface of radius RR is drawn as a full circle centered on the optical axis. A ray meets the surface at height cc above the axis. The slanted radius from the center to that hit point makes an angle θS\theta_S with the horizontal axis-radius: and the same angle reappears at the surface, between the vertical reference direction and the surface normal. In the small triangle formed by the slanted radius, the horizontal axis-radius, and the vertical leg of height cc, basic trigonometry gives sin⁡θS=c/R\sin\theta_S = c/R. In the paraxial regime, this simplifies to θS≈cR,\theta_S \approx \frac{c}{R}, which is exactly what the surface_angle helper above returns, and what the lensmaker derivation in Figure 6.4(b) used at each surface. (Book Figure 6.5.) Output from cell 11 With θS≈c/R\theta_S \approx c/R established, the surface-tilt relation above can drop cc on both sides: completing the cancellation that lets one image distance bb serve every ray from the same object point.

Off-axis points

The derivation so far placed the object on the optical axis. Off-axis points need only a small extension: rotating the whole construction through a small angle θR\theta_R adds θR\theta_R to θS\theta_S at each surface in Table 6.1, leaving the algebra unchanged. The lensmaker’s formula still holds, and the image lands at the conjugate distance bb on the far side, at height P2=−b P1/aP_2 = -b\,P_1 / a: inverted, below the axis. In the paraxial, thin-lens limit, this promotes the focusing property from points on an axis to whole planes: every point on the object plane at distance aa focuses to a corresponding point on the image plane at distance bb, both perpendicular to the optical axis. Figure 6.6: Off-axis points and the equivalent lens rotation. The off-axis object point P1P_1 sits at height P1P_1 above the optical axis, at distance aa from the lens; P0P_0 marks the on-axis reference. Two rays trace the image: a parallel ray refracts through the focal point at distance ff, and a ray through the lens center proceeds undeviated. They meet at the image point on the image plane at distance bb. The angle θR\theta_R marks the small rotation that makes the off-axis case equivalent to the on-axis derivation. (Book Figure 6.6.) Output from cell 12 With Figure 6.6, the derivation is complete: the lensmaker’s formula 1a+1b=1f,1f=(n−1)(1R1+1R2),\frac{1}{a} + \frac{1}{b} = \frac{1}{f}, \qquad \frac{1}{f} = (n-1)\left(\frac{1}{R_1} + \frac{1}{R_2}\right), describes how a thin lens maps every point on an object plane at distance aa to a corresponding point on the image plane at distance bb. The book’s Figure 6.7, a photograph of a laser pointer swept across a lens with every ray converging to the same spot on the wall, demonstrates the same focusing property physically. The next section takes the lensmaker’s formula and puts it to work: straight-through rays at the lens center, conjugate-point ray tracing, depth of field, concave lenses, and a Galilean telescope built from two of them.

Imaging with lenses

With the lensmaker’s formula in hand, the rest of the chapter puts it to work: predicting where images form, how sharp they are, what changes with a concave lens, and how two lenses combine into a telescope. The starting observation is a small one, but it’s what lets the whole apparatus mimic a pinhole camera. At the very center of a thin lens, the front and back surfaces are parallel. A ray crossing that region refracts at the first surface, then refracts back through the same angle at the second surface: emerging parallel to the way it entered, displaced laterally by a small amount that depends on the lens’s thickness. In the thin-lens limit, the two surfaces collapse to a single plane: the displacement vanishes and the ray passes straight through. Every direction through the center is undeviated, which means the lens behaves, for those center rays, exactly like a pinhole: and a thin lens therefore images the world in perspective projection, just as a pinhole does. Figure 6.8: Rays through the center of a thin lens. (a) A physical lens of non-zero thickness: the ray refracts at both surfaces and exits parallel to the incoming direction, with a small lateral displacement. (b) The same ray under the thin-lens approximation: the two surfaces collapse to a single plane, the displacement vanishes, and the ray passes straight through. (c) Because every direction through the center is undeviated, a fan of rays through the lens center behaves identically to a fan of rays through a pinhole. (Book Figure 6.8.) Output from cell 13 With this observation in place, four properties now characterize how rays travel through a thin lens:
  1. Focusing: every ray from a point at distance aa reconverges at the conjugate distance bb, with 1a+1b=1f\frac{1}{a} + \frac{1}{b} = \frac{1}{f}.
  2. Parallel rays: the limit a→∞a \to \infty: parallel rays converge at the focal point, distance ff behind the lens.
  3. Center rays: any ray through the lens center proceeds in a straight line, as panel (c) above shows.
  4. Magnification: combining the first and third, a plane at distance aa images with scale b/ab/a: exactly as a pinhole at the lens center would project it.
Two points on opposite sides of the lens at distances aa and bb that satisfy the lensmaker’s formula are called conjugate points: light from one focuses to the other, and vice versa. The next figure traces conjugate-point pairs for several object distances, showing how the image distance moves as the object slides toward or away from the lens. The four ray-tracing rules above pin down the image distance for any object distance through 1a+1b=1f\frac{1}{a} + \frac{1}{b} = \frac{1}{f}. The two endpoints of the formula are limits: object at infinity, image at ff; object at ff, image at infinity: and everything in between trades smoothly between them. The next figure traces five cases on the same lens, holding the focal length fixed and moving the object closer step by step. Figure 6.9: Conjugate points for a convex thin lens. All five panels show the same lens with focal points marked at ±f\pm f from the lens (cyan dots). (a) Parallel rays from infinity converge at the right-side focal point: the limiting case a→∞a \to \infty, b=fb = f. (b) Object at a=3fa = 3f images at b=1.5fb = 1.5f. (c) Object at a=2fa = 2f images symmetrically at b=2fb = 2f: the unit-magnification case. (d) Object at a=1.5fa = 1.5f images at b=3fb = 3f: as the object moves inward, the image moves outward. (e) Object at the left focal point (a=fa = f) produces parallel rays on exit: the other limiting case, b→∞b \to \infty. (Book Figure 6.9.) Output from cell 14 Set the lens up inside a camera. The sensor sits at a fixed distance bb behind the lens; the object lives somewhere out in the world at distance aa. If aa and bb happen to satisfy the lensmaker’s formula, the object is in focus: every ray from a surface point converges to a single point on the sensor, and the image is sharp. But aa varies with what the camera is pointed at, and bb is fixed by the camera’s geometry. When aa doesn’t quite match the conjugate of bb, rays from a single object point don’t converge at the sensor: they hit it as a small circle of confusion. How small can that circle get before the image looks blurry, and how much of the scene stays acceptably sharp at once? That’s the depth of field, and it’s what comes next.

Depth of field

When the lens lives inside a camera, the sensor sits at a fixed distance behind it, so only one object distance is in sharp focus: the one whose conjugate is the sensor itself. That object distance defines an object plane (the focal plane in the book’s terminology) where everything images sharply. Objects in front of or behind that plane image past or before the sensor, hitting it as a small disk: a circle of confusion. A photograph still looks sharp as long as that circle is small enough that the eye can’t tell, and the range of object distances over which it stays acceptable is the depth of field. Figure 6.10: Circle of confusion and depth of field. Three object positions, one fixed sensor plane. The top row shows an object on the in-focus plane: every ray converges to a single point on the sensor for a sharp image. The middle row shows an object closer than the in-focus plane: its image forms past the sensor, hitting the sensor as a circle of confusion. The bottom row shows an object further than the in-focus plane: its image forms before the sensor, again hitting as a circle of confusion. The bracket on the left marks the depth of field. (Book Figure 6.10.) Output from cell 15 The depth of field is set by the lens’s focal length, the maximum circle-of-confusion size that still looks sharp, and the camera’s f-number N=f/AN = f / A, where AA is the aperture diameter. The next figure lays out the variables. The geometry that turns the circle of confusion into a depth-of-field range is two pairs of similar triangles: one for the near limit, one for the far. Figure 6.11: Variables for the depth-of-field calculation. Two scenes, same focal length, same focused object distance UU, same circle of confusion CC at the sensor: only the aperture differs. The f-number N=f/AN = f/A relates aperture diameter AA to focal length: smaller NN means a wider aperture. (Top, NaN^a) Wider aperture, narrower depth of field: the range D1aD_1^a to D2aD_2^a around UU is short. (Bottom, NbN^b) Narrower aperture, wider depth of field: same CC at the sensor, but a much larger range D1bD_1^b to D2bD_2^b around UU. (Book Figure 6.11.) Output from cell 16 The two similar-triangle pairs in Figure 6.11 yield expressions for D1D_1 and D2D_2 that, when solved and summed, give the exact depth of field: D=2NCU2f2f4−N2C2U2.D = \frac{2 N C U^2 f^2}{f^4 - N^2 C^2 U^2}. For the practical regime C≪f/NC \ll f/N, the second term in the denominator drops out and the formula collapses to a clean proportionality: D≈2NCU2f2.D \approx \frac{2 N C U^2}{f^2}. Depth of field is linear in NN. Doubling NN doubles the in-focus range: but the light hitting the sensor falls as 1/N21/N^2, so the same scene needs four times the exposure. The next figure shows the trade-off on a real (synthetic) ruler. Figure 6.12: Photographic depth of field as a function of aperture. One sharp photograph of a ruler, blurred at each pixel by an amount proportional to its depth offset from the focal plane and scaled by 1/N1/N: recreating the book’s Figure 6.12 photographic demonstration. (a) f/2f/2: narrow sharp band, blurred ends. (b) f/4f/4: doubled DOF. (c) f/8f/8: nearly the whole ruler sharp. (Book Figure 6.12. Recreated by applying a Gaussian blur with σ∝1/N\sigma \propto 1/N; the f/2f/2 reference blur is calibrated visually.) Output from cell 17 The photographic trade-off, wider aperture, more light but shallower DOF, is what the formula D≈2NCU2/f2D \approx 2NCU^2/f^2 makes quantitative. The next part turns the lens curvature inside out and asks what changes with a concave lens.

Concave lenses

The lens designed above was convex: both surfaces bowing outward, focal length positive. Reverse the curvature on both surfaces and you get a concave lens, with both surfaces bowing inward. In the lensmaker’s formula 1f=(n−1)(1R1+1R2),\frac{1}{f} = (n-1)\left(\frac{1}{R_1} + \frac{1}{R_2}\right), a concave surface contributes the same magnitude with the opposite sign, so ff comes out negative. The same paraxial geometry that focuses rays to a real point past a convex lens now bends them away from the axis at a concave lens: but the back-projection of those diverging rays still meets at a single point, on the source side of the lens. That point is the lens’s virtual focal point. Figure 6.13: Convex and concave thin-lens behavior. (a) A convex lens with focal length +f+f: parallel rays from the left converge to a real focal point at distance ff past the lens. (b) A concave lens with focal length −f-f: parallel rays diverge after the lens, but their back-projections (dotted) meet at a virtual focal point at distance ff on the source side. (c) Same concave lens with a tilted parallel bundle: the virtual focal point shifts off-axis (cyan), just as the focal point in a convex lens shifts to image off-axis sources: the lensmaker’s formula handles both cases with one sign change. (Book Figure 6.13.) Output from cell 18 A concave lens alone can’t form a real image: it diverges every bundle it sees. But paired with a convex lens, that diverging behavior becomes useful: if the concave lens sits at the convex lens’s focal point, the two together turn a converging bundle back into parallel rays: at a different angle than they entered. That angular amplification is the principle behind the Galilean telescope, which comes next.

Lenses in a telescope

A convex lens and a concave lens together form a Galilean telescope, named after the one Galileo built in 1609. The construction is simple: place a convex lens (lens 1, focal length f1f_1, long) with a concave lens (lens 2, focal length f2f_2, shorter) such that they share a common focal point: lens 1 is distance f1f_1 to the left of that shared point, lens 2 is distance f2f_2 to the same side. The lenses sit f1−f2f_1 - f_2 apart. In that configuration, parallel rays entering lens 1 converge toward the shared focal point, and lens 2, intercepting them before they reach it, refracts them back into a parallel bundle. The output bundle is parallel like the input, but compressed into a smaller cross-section. What makes it a telescope is what happens when the input direction tilts. Figure 6.14: Galilean telescope geometry. (a) Parallel input, parallel output. Three parallel rays enter lens 1 from the left, converge toward the shared focal point, and are intercepted by lens 2 before reaching it: refracting back into a parallel bundle on the right. The output rays are parallel like the input, but closer together, compressed into the smaller exit aperture. (b) Tilted input, angular magnification. When the input bundle tilts at angle δi\delta_i from the optical axis, the chief ray through lens 1’s center reaches the shared focal point at height d=f1δid = f_1 \delta_i above the axis (small-angle approximation, property 3 above). The same point dd is at distance f2f_2 from lens 2, which refracts the bundle into a parallel output at angle δo=d/f2\delta_o = d / f_2 from the axis. The two similar triangles share the height dd but have different base lengths f1f_1 and f2f_2, giving the magnification M=δo/δi=f1/f2M = \delta_o / \delta_i = f_1 / f_2. (Book Figure 6.14.) Output from cell 19 With f1>f2f_1 > f_2 the telescope magnifies. Galileo built his at M≈30M \approx 30 by pairing a long convex objective with a short concave eyepiece; the book’s Figure 6.15 and Figure 6.16 show a cardboard recreation (M≈27M \approx 27 from f1=500f_1 = 500 mm, f2=18f_2 = 18 mm) and the moon through it, alongside Galileo’s own lunar drawings. With this, the imaging part is complete. This section has worked from Snell’s law through the lensmaker’s formula and used it for imaging, depth of field, concave lenses, and the telescope.