Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #73), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 47 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. When a 3D point moves, or the camera moves, how does its image on the camera plane move? This section answers that question with one long derivation: the motion field equation (47.8). Once you have it, every interesting motion phenomenon falls out as a special case: vanishing points, parallax, the focus of expansion, time-to-contact, and the way a zoom looks different from a dolly. Each figure here is a diagram or a sketch of a vector field, drawn by encoding the underlying math in PyTorch. There are no photographs.

Perspective projection of a moving point

A 3D point P=(X,Y,Z)\mathbf{P} = (X, Y, Z) projects to image coordinates p=(x,y)\mathbf{p} = (x, y) through the pinhole camera with focal length ff (equation 47.9): x=f XZ,y=f YZ.x = f\,\frac{X}{Z}, \qquad y = f\,\frac{Y}{Z}. When P\mathbf{P} moves with velocity P˙=(X˙,Y˙,Z˙)\dot{\mathbf{P}} = (\dot X, \dot Y, \dot Z), the image point moves too. Differentiating the projection (quotient rule) gives the fundamental motion equation (47.1): [x˙y˙]=1Z[f0−x0f−y][X˙Y˙Z˙].\begin{bmatrix}\dot x\\\dot y\end{bmatrix} = \frac{1}{Z}\begin{bmatrix} f & 0 & -x \\ 0 & f & -y \end{bmatrix}\begin{bmatrix}\dot X\\\dot Y\\\dot Z\end{bmatrix}. The 1/Z1/Z factor is the entire reason vision can recover depth from motion, because closer points produce larger image-plane velocities than farther points moving the same way.
Output from cell 4

Two easy cases of equation 47.1

Equation 47.1 has three terms. Killing one at a time isolates the two simplest motion regimes: Motion parallel to the image plane (Z˙=0\dot Z = 0, equation 47.2): [x˙y˙]=fZ[X˙Y˙].\begin{bmatrix}\dot x\\\dot y\end{bmatrix} = \frac{f}{Z}\begin{bmatrix}\dot X\\\dot Y\end{bmatrix}. The magnitude is inversely proportional to depth: pure parallax. Motion along the optical axis (X˙=Y˙=0\dot X = \dot Y = 0, equation 47.3): [x˙y˙]=−Z˙Z[xy].\begin{bmatrix}\dot x\\\dot y\end{bmatrix} = -\frac{\dot Z}{Z}\begin{bmatrix}x\\y\end{bmatrix}. Motion is radial, scaled by the ratio Z˙/Z\dot Z / Z (the time-to-contact). Figure 47.2 sketches both. Output from cell 5

The vanishing point

Let a point move with constant world velocity V=(VX,VY,VZ)\mathbf{V} = (V_X, V_Y, V_Z) from initial position P0\mathbf{P}_0. Its position is P(t)=P0+tV\mathbf{P}(t) = \mathbf{P}_0 + t\mathbf{V}, and its image is p(t)=fP0,Z+tVZ[P0,X+tVXP0,Y+tVY].\mathbf{p}(t) = \frac{f}{P_{0,Z} + tV_Z}\begin{bmatrix}P_{0,X} + tV_X \\ P_{0,Y} + tV_Y\end{bmatrix}. As t→∞t \to \infty the initial-position terms drop out and you land at the vanishing point (equation 47.4): x∞=f VXVZ,y∞=f VYVZ.x_{\infty} = f\,\frac{V_X}{V_Z},\qquad y_{\infty} = f\,\frac{V_Y}{V_Z}. The vanishing point depends only on the direction of motion, never on where the point started. Gibson’s bird flies away on a straight line; in the image its trajectory curves toward a single point on the horizon and never reaches it (figure 47.3). Output from cell 6

The motion field under camera translation

When the camera translates with velocity V\mathbf{V} through a static world, each scene point moves with −V-\mathbf{V} in the camera frame. Substituting P˙=−V\dot{\mathbf{P}} = -\mathbf{V} into equation 47.1 gives the camera-translation motion field (equation 47.5): [x˙y˙]=1Z[−f0x0−fy][VXVYVZ].\begin{bmatrix}\dot x\\\dot y\end{bmatrix} = \frac{1}{Z}\begin{bmatrix}-f & 0 & x\\0 & -f & y\end{bmatrix}\begin{bmatrix}V_X\\V_Y\\V_Z\end{bmatrix}. The full motion field including rotation is given by equation 47.8, which the next cell encodes once and every later figure reuses.

Lateral camera motion

Driving past a roadside scene: VX≠0V_X \ne 0, VY=VZ=0V_Y = V_Z = 0. The motion field collapses to x˙=−fVXZ,y˙=0.\dot x = -\frac{f V_X}{Z}, \qquad \dot y = 0. Image-plane speed is inversely proportional to scene depth: close objects whip past while distant clouds barely move (figure 47.5). Output from cell 8 Output from cell 9

Forward camera motion and the focus of expansion

Driving toward a wall: VX=VY=0V_X = V_Y = 0, VZ>0V_Z > 0. Equation 47.5 simplifies to x˙=VZZ x,y˙=VZZ y,\dot x = \frac{V_Z}{Z}\,x, \qquad \dot y = \frac{V_Z}{Z}\,y, a radial flow centered at the origin. The point at which the field is zero is the focus of expansion, and the scaling factor VZ/ZV_Z / Z is the inverse time-to-contact, the quantity a fly’s-eye control loop reads off to time a landing. Output from cell 10 Output from cell 11

Camera rotation and the full motion field

Adding rotation, the point’s velocity relative to a rotating camera with angular velocity Ω\mathbf{\Omega} is P˙=−V−Ω×P.\dot{\mathbf{P}} = -\mathbf{V} - \mathbf{\Omega} \times \mathbf{P}. Plugging this into equation 47.1 produces the full motion field (equation 47.8): [x˙y˙]=1Z[−f0x0−fy][VXVYVZ]⏟translation: depends on Z  +  1f[xy−(f2+x2)fyf2+y2−xy−fx][ΩXΩYΩZ]⏟rotation: independent of Z.\begin{bmatrix}\dot x\\\dot y\end{bmatrix} = \underbrace{\frac{1}{Z}\begin{bmatrix}-f & 0 & x\\0 & -f & y\end{bmatrix}\begin{bmatrix}V_X\\V_Y\\V_Z\end{bmatrix}}_{\text{translation: depends on }Z} \;+\; \underbrace{\frac{1}{f}\begin{bmatrix}xy & -(f^2+x^2) & fy\\ f^2+y^2 & -xy & -fx\end{bmatrix}\begin{bmatrix}\Omega_X\\\Omega_Y\\\Omega_Z\end{bmatrix}}_{\text{rotation: independent of }Z}. Two qualitative consequences are worth noting:
  • Rotational flow is independent of depth, so it tells you nothing about the scene’s 3D structure.
  • Translational flow scales as 1/Z1/Z, so it carries all the depth information.
Figure 47.8 introduces the three rotation axes (yaw, pitch, roll) used throughout the rest of the chapter. Output from cell 12

Rotation around the optical axis (ΩX=ΩY=0\Omega_X = \Omega_Y = 0)

With only ΩZ\Omega_Z active, equation 47.8 collapses to x˙=ΩZ y,y˙=−ΩZ x,\dot x = \Omega_Z\,y,\qquad \dot y = -\Omega_Z\,x, the velocity field of rigid rotation about the image origin. Every concentric circle is an integral curve (figure 47.9). Output from cell 13

Rotation under varying focal length

Y-axis rotation with ΩX=ΩZ=0\Omega_X = \Omega_Z = 0 produces x˙=−(f2+x2)f ΩY,y˙=−xyf ΩY,\dot x = -\frac{(f^2 + x^2)}{f}\,\Omega_Y,\qquad \dot y = -\frac{xy}{f}\,\Omega_Y, which depends strongly on ff. Wide-angle lenses (ff small) produce flows with sharp curvature near the edges; long lenses (ff large) produce nearly uniform horizontal flow. Figure 47.10 sweeps f∈{1/3,1,3}f \in \{1/3, 1, 3\} with the same scene and angular velocity. Output from cell 14

Concluding remarks

The whole chapter rests on three ideas:
  1. Differentiate the projection p=(fX/Z,fY/Z)\mathbf{p} = (fX/Z, fY/Z) to get the image-plane motion of any moving 3D point (eq. 47.1).
  2. Substitute the camera’s rigid-body kinematics P˙=−V−Ω×P\dot{\mathbf{P}} = -\mathbf{V} - \mathbf{\Omega}\times\mathbf{P} to get the motion field (eq. 47.8).
  3. Read off special cases: vanishing points (47.4), depth-modulated parallax (47.5), focus of expansion and time-to-contact (47.6), depth-independent rotation (47.7), focal-length sensitivity (47.10).
The next chapter of the book (48, optical flow) flips the script: given two frames, recover the motion field. The geometry here is the forward model that chapter inverts.