Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #74), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 38 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. This section lays the foundation for all the geometry that follows. Three ideas drive it:
  1. Homogeneous coordinates let you write translation, rotation, scaling, shearing, and projective warps as a single matrix multiplication.
  2. Lines and points are dual under the cross product in homogeneous form: joining two points and intersecting two lines are the same operation.
  3. Implicit image representations (SIREN-style networks that map (x,y)→intensity(x,y)\to \text{intensity}) give you sub-pixel image access without an interpolation kernel, so geometric transformations become inverse-coordinate evaluations.
The figures are regenerated in PyTorch. Two of them go furthest: figure 38.8 reframes warping as a 1D convolution on coordinates, and figure 38.12 replaces the pixel grid with a neural network.

Homogeneous and heterogeneous coordinates

Heterogeneous coordinates write a 2D point as (x,y)(x, y). Homogeneous coordinates write the same point as (x,y,1)(x, y, 1), and also as (λx,λy,λ)(\lambda x, \lambda y, \lambda) for any λ≠0\lambda \ne 0. All points on the ray through the origin and (x,y,1)(x, y, 1) represent the same 2D point. Converting back means dividing by ww: [xyw]  ⟶  (xw,yw).\begin{bmatrix}x\\y\\w\end{bmatrix} \;\longrightarrow\; \left(\frac{x}{w}, \frac{y}{w}\right). The payoff is that translation, an addition in heterogeneous coordinates, becomes a multiplication in homogeneous coordinates. This uniformity lets you compose any sequence of geometric transformations into a single matrix product.
Output from cell 5

2D image transformations

Every transformation that follows is a 3×33\times 3 matrix acting on homogeneous points. The book applies each transformation to a photograph of a clock. Here a wall-clock image stands in for it, and each matrix is applied as an actual image warp with kornia.geometry.transform.warp_affine, so what you see is the matrix acting on real pixels.
Output from cell 10 Output from cell 11 Output from cell 12 Output from cell 13 Output from cell 14 Output from cell 15

Geometric transformation as a convolution

If you represent an image as a set of triples {(ℓi,xi,yi)}\{(\ell_i, x_i, y_i)\} (intensity plus explicit position), then a rotation is a 1D convolution on the position channels: [x′y′]=[cos⁡θ−sin⁡θsin⁡θcos⁡θ]⏟kernels wx,wy[xy].\begin{bmatrix}x'\\y'\end{bmatrix} = \underbrace{\begin{bmatrix}\cos\theta & -\sin\theta\\\sin\theta & \cos\theta\end{bmatrix}}_{\text{kernels } w_x, w_y} \begin{bmatrix}x\\y\end{bmatrix}. Intensity passes through untouched; only the coordinate channels are mixed (figure 38.8). This framing makes it natural to learn warps with neural networks that operate on coordinate channels, and the rest of the section builds that bridge. Output from cell 16

Lines and points: cross-product duality

In homogeneous coordinates, a 2D line ax+by+c=0ax + by + c = 0 is the vector l=(a,b,c)\mathbf{l} = (a, b, c), and incidence is l⊤p=0\mathbf{l}^\top \mathbf{p} = 0. Two operations turn out to be the same cross product: l=p1×p2(line through two points)\mathbf{l} = \mathbf{p}_1 \times \mathbf{p}_2 \qquad \text{(line through two points)} p=l1×l2(intersection of two lines)\mathbf{p} = \mathbf{l}_1 \times \mathbf{l}_2 \qquad \text{(intersection of two lines)} Parallel lines intersect at a point with w=0w = 0: a point at infinity, the homogeneous form of a vanishing point. The section on 3D motion and its 2D projection uses exactly this construction. Output from cell 17

Image warping: forward and backward mapping

Applying a transformation MM to an image should give you a new image. The naive way, forward mapping, walks the input pixels, transforms each location, and writes its intensity to the nearest target pixel. That leaves holes, because an expansion spreads the input pixels too thinly to cover the target grid. The correct way, backward mapping, walks the target pixels and pulls each back through M−1M^{-1}, so every output pixel gets one well-defined intensity (figures 38.10 and 38.11). A small checkerboard makes the artifacts easy to see.
Output from cell 19 Output from cell 19

Differentiable warping with Kornia

The hand-written backward_warp above shows the principle. In practice you use Kornia, a differentiable computer vision library built on PyTorch. Its warp_affine performs the same backward mapping, but with bilinear interpolation (so no nearest-neighbor aliasing), in batches, and differentiably end to end. The next cell feeds it the same matrix MM and compares the result with the hand-written version.
Output from cell 21

Implicit image representations (SIREN)

Backward mapping at sub-pixel locations needs an interpolation kernel, but only because the image is stored as a discrete grid. SIREN (Sitzmann et al., NeurIPS 2020) does away with the grid: it represents an image as a small network fθ:(x,y)→intensityf_\theta : (x, y) \to \text{intensity} whose layers use sin⁡\sin activations. The trained network is the image, and you can evaluate it at any continuous coordinate. Geometric warping then becomes a coordinate transform with no interpolation kernel: ℓ^(x,y)=fθ(M−1(x,y,1)⊤).\hat\ell(x, y) = f_\theta(M^{-1}(x, y, 1)^\top). The next cells train a small SIREN on the cameraman test image and warp it by feeding pre-transformed coordinates (figure 38.12). The training is deliberately short: the point is to demonstrate the representation, not to maximize reconstruction quality.
Output from cell 24

Concluding remarks

The section moves forward one promotion at a time:
  • Heterogeneous to homogeneous: every affine and projective transformation becomes a single matrix product.
  • Points and lines to cross-product duality: incidence, joins, and intersections collapse into one operation. Parallel lines meet at points with w=0w=0, so vanishing points get a coordinate.
  • Discrete grid to implicit function: with SIREN, the image is its parameters. Warping becomes an inverse-coordinate query and interpolation kernels disappear.
Every later geometry chapter of the book (perspective projection in 39, stereo in 40, homographies in 41, single-view metrology in 42, structure from motion in 44, radiance fields in 45) reuses the homogeneous algebra and the point and line duality set up here.