Skip to main content
Open In Colab This section was written by Kimberly Milner (pull request #51), with help from an AI coding agent on the code. It reproduces the ideas of Chapter 15 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. Image filtering is the workhorse of low-level vision: blurring, sharpening, edge detection, and template matching are all the same operation, the convolution of an image with a small kernel. This section regenerates the figures in PyTorch and Kornia, building from continuous and discrete signals, through linear translation-invariant (LTI) systems and convolution, to cross-correlation and template matching. The thread through the section. A system maps an input signal to an output. The systems worth studying are linear and translation-invariant, and every LTI system is completely described by a single impulse response hh, applied everywhere by convolution: ℓout=ℓin∗h\ell_\text{out} = \ell_\text{in} * h. Cross-correlation is the same computation without flipping the kernel, which is what template matching uses. Where the book shows a photograph or a hand sketch, a link points to it, and the filtering operations run on scikit-image test images instead.
Images replaced for licensing. The book is published under a CC BY-NC-ND license, which covers only the book as a whole and not its individual images, so this page does not republish the book’s photographs. License-free images stand in for them:If you use the book for non-commercial purposes, you can swap the originals back in: each line that loads a stand-in carries the original’s link in a comment.

Signals and images

A one-dimensional continuous signal ℓ(t)\ell(t) is defined for every real tt; sampling it on an integer grid gives the discrete signal ℓ[n]\ell[n]. Images are just 2-D discrete signals ℓ[n,m]\ell[n,m]. Figure 15.1: a continuous signal and its discrete samples.
Book Figure 15.1: computed here.
Output from cell 5

Systems

A system ff maps an input signal to an output signal. The systems that matter here are linear (they respect scaling and superposition) and translation-invariant (shifting the input just shifts the output). Figure 15.2: a system maps an input signal to an output.
Book Figure 15.2: drawn as a schematic.
Output from cell 6 Figure 15.3: which of these image transforms are linear?.
Book Figure 15.3: applied with Kornia.
Output from cell 8 Figure 15.4: dense vs. local connectivity. A general linear filter is a matrix multiply y=Hxy = Hx in which every output pixel depends on every input pixel (a fully-connected layer). Convolution restricts HH to local connections: each output uses only a small window of nearby inputs, which is what makes filtering practical. (Book Figure 15.4: computed here.)
Output from cell 10

Convolution

An LTI system applies one impulse response hh at every position: the output is the input convolved with hh. Convolution flips the kernel and slides it; it is linear, translation-invariant, commutative and associative. Figure 15.5: translation invariance. An object looks the same wherever it appears in the frame, so a useful filter must respond identically at every position: shifting the input just shifts the output, f(shift(x))=shift(f(x))f(\text{shift}(x)) = \text{shift}(f(x)). The book shows this with a photograph of birds; here the same blur runs on a scikit-image photograph.
Output from cell 12 Figure 15.6: convolution is weight sharing. Where 15.4’s dense operator wires each output to every input, a convolution wires each output to a small local window using the same kernel weights at every position. That one shared, shifted kernel is exactly what a banded Toeplitz matrix encodes, and it is what makes the filter translation-invariant. (Book Figure 15.6: computed here.)
Output from cell 14 Figure 15.7: 2-D convolution of a 9×9 image with a 3×3 kernel, step by step.
Book Figure 15.7: computed here.
Output from cell 16 Figure 15.8: defocus is a convolution; rotation is not.
Book Figure 15.8: applied with Kornia.
Output from cell 18 Figure 15.9: four convolutions: identity, shift, superposition, averaging.
Book Figure 15.9: applied with Kornia.
Output from cell 20 Figure 15.10: handling boundaries: zero, circular, mirror and replicate padding.
Book Figure 15.10: applied with Kornia.
Output from cell 22

Cross-correlation versus convolution

Cross-correlation is convolution without flipping the kernel. For template matching you slide a template over the image and measure similarity; normalized cross-correlation removes the effect of local brightness. Figure 15.11: cross-correlation vs. convolution (kernel orientation).
Book Figure 15.11: computed here.
Output from cell 24 Figure 15.12: template matching for the letter “a” via normalized cross-correlation.
Book Figure 15.12: computed here.
Output from cell 26

System identification

An LTI system is identified by its impulse response: feed in an impulse δ[n]\delta[n] and read out h[n]h[n]. A room’s acoustics are an LTI system: a clap measures its impulse response, and any sound is that sound convolved with hh. Figure 15.13: an LTI system is fully described by its impulse response h[n]h[n].
Book Figure 15.13: drawn as a schematic.
Output from cell 27 Figure 15.14: room acoustics. Sound reaches a microphone along many paths, a direct one and many reflections. The book shows this with a sketch of a room. Figure 15.15: measuring a room’s impulse response. A short, loud sound such as a clap is close to an impulse, so what the microphone records is the room’s impulse response. The book shows the measurement setup. Figure 15.16: a voice signal convolved with a room impulse response.
Book Figure 15.16: computed here.
Output from cell 29

Concluding remarks

Every operation in this section is one idea in different forms. A linear translation-invariant system is a matrix multiply whose matrix is banded and Toeplitz (Figures 15.4 and 15.6). Equivalently, it is a convolution with a single kernel, the impulse response, applied at every position (Figures 15.6 to 15.9). Blurring, shifting, averaging, edge detection, and template matching are all just different kernels. Cross-correlation is the same computation without the flip (Figure 15.11), and with local normalization it becomes a robust template detector (Figure 15.12). Because an LTI system is fully described by its impulse response, you can identify it by measuring that response, whether the system is a lens, a filter, or a room (Figures 15.13 to 15.16).