- Homogeneous coordinates let you write translation, rotation, scaling, shearing, and projective warps as a single matrix multiplication.
- Lines and points are dual under the cross product in homogeneous form: joining two points and intersecting two lines are the same operation.
- Implicit image representations (SIREN-style networks that map ) give you sub-pixel image access without an interpolation kernel, so geometric transformations become inverse-coordinate evaluations.
Homogeneous and heterogeneous coordinates
Heterogeneous coordinates write a 2D point as . Homogeneous coordinates write the same point as , and also as for any . All points on the ray through the origin and represent the same 2D point. Converting back means dividing by : The payoff is that translation, an addition in heterogeneous coordinates, becomes a multiplication in homogeneous coordinates. This uniformity lets you compose any sequence of geometric transformations into a single matrix product.
2D image transformations
Every transformation that follows is a matrix acting on homogeneous points. The book applies each transformation to a photograph of a clock. Here a wall-clock image stands in for it, and each matrix is applied as an actual image warp withkornia.geometry.transform.warp_affine, so what you see is the matrix acting on real pixels.






Geometric transformation as a convolution
If you represent an image as a set of triples (intensity plus explicit position), then a rotation is a 1D convolution on the position channels: Intensity passes through untouched; only the coordinate channels are mixed (figure 38.8). This framing makes it natural to learn warps with neural networks that operate on coordinate channels, and the rest of the section builds that bridge.
Lines and points: cross-product duality
In homogeneous coordinates, a 2D line is the vector , and incidence is . Two operations turn out to be the same cross product: Parallel lines intersect at a point with : a point at infinity, the homogeneous form of a vanishing point. The section on 3D motion and its 2D projection uses exactly this construction.
Image warping: forward and backward mapping
Applying a transformation to an image should give you a new image. The naive way, forward mapping, walks the input pixels, transforms each location, and writes its intensity to the nearest target pixel. That leaves holes, because an expansion spreads the input pixels too thinly to cover the target grid. The correct way, backward mapping, walks the target pixels and pulls each back through , so every output pixel gets one well-defined intensity (figures 38.10 and 38.11). A small checkerboard makes the artifacts easy to see.

Differentiable warping with Kornia
The hand-writtenbackward_warp above shows the principle. In practice you use Kornia, a differentiable computer vision library built on PyTorch. Its warp_affine performs the same backward mapping, but with bilinear interpolation (so no nearest-neighbor aliasing), in batches, and differentiably end to end. The next cell feeds it the same matrix and compares the result with the hand-written version.

Implicit image representations (SIREN)
Backward mapping at sub-pixel locations needs an interpolation kernel, but only because the image is stored as a discrete grid. SIREN (Sitzmann et al., NeurIPS 2020) does away with the grid: it represents an image as a small network whose layers use activations. The trained network is the image, and you can evaluate it at any continuous coordinate. Geometric warping then becomes a coordinate transform with no interpolation kernel: The next cells train a small SIREN on the cameraman test image and warp it by feeding pre-transformed coordinates (figure 38.12). The training is deliberately short: the point is to demonstrate the representation, not to maximize reconstruction quality.
Concluding remarks
The section moves forward one promotion at a time:- Heterogeneous to homogeneous: every affine and projective transformation becomes a single matrix product.
- Points and lines to cross-product duality: incidence, joins, and intersections collapse into one operation. Parallel lines meet at points with , so vanishing points get a coordinate.
- Discrete grid to implicit function: with SIREN, the image is its parameters. Warping becomes an inverse-coordinate query and interpolation kernels disappear.

