Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #76), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 17 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. Blur filters are low-pass linear filters: they attenuate high spatial frequencies (fine detail and noise) while preserving the low-frequency structure of an image. This section builds the three filter families of the book’s chapter, the box filter, the Gaussian filter, and the binomial filter, implements each as a convolution, and regenerates the figures from the underlying math. Four of the book’s figures (17.1, 17.4, 17.5, and 17.8) are built on its own photographs. Here the same experiments run on scikit-image test images, and each links to the book’s version.
Images replaced for licensing. The book is published under a CC BY-NC-ND license, which covers only the book as a whole and not its individual images, so this page does not republish the book’s photographs. License-free images stand in for them:If you use the book for non-commercial purposes, you can swap the originals back in: each line that loads a stand-in carries the original’s link in a comment.

Noise removal versus detail loss

A blur filter replaces each pixel with a weighted average of its neighbors. Averaging suppresses zero-mean noise, because the fluctuations cancel, but it also smears genuine high-frequency detail. That is the central tradeoff of the whole chapter. Figure 17.1 makes it visible: additive noise is largely gone after a 5×55\times5 average, at the cost of sharpness. The book uses a noisy photograph of a stop sign; here noise is added to a color test photograph.
Output from cell 5

The box filter

The 2D box (moving-average) kernel keeps a constant weight inside a rectangular window and zero outside: boxN,M[n,m]={1−N≤n≤N, −M≤m≤M0otherwise\mathrm{box}_{N,M}[n,m] = \begin{cases}1 & -N \le n \le N,\ -M \le m \le M\\ 0 & \text{otherwise}\end{cases} To keep average brightness unchanged the kernel must have DC gain 1, i.e. its coefficients sum to 1, so you divide by (2N+1)(2M+1)(2N+1)(2M+1). The box is separable: boxN,M=boxN x∗boxM y\mathrm{box}_{N,M} = \mathrm{box}_N^{\,x} * \mathrm{box}_M^{\,y}, a full-window rectangle equals a horizontal bar convolved with a vertical bar. Choosing N=0N=0 or M=0M=0 blurs along a single axis, as the middle and right panels of Figure 17.2 show.
Output from cell 8

The box filter’s frequency response is not monotonic

The discrete-time Fourier transform of a length-LL box is a Dirichlet (aliased-sinc) kernel. Because a sinc oscillates, the box’s frequency response has side lobes: some high frequencies are passed with more gain than lower ones, and the sign flips lobe-to-lobe. A good low-pass filter should instead fall off monotonically. That is the motivation for the Gaussian and binomial filters below.
Output from cell 11

The Gaussian filter

The Gaussian is the canonical isotropic blur. Continuous form: g(x;σ)=12πσ2 e−x2/(2σ2),g(x,y;σ)=12πσ2 e−(x2+y2)/(2σ2).g(x;\sigma) = \frac{1}{\sqrt{2\pi\sigma^2}}\,e^{-x^2/(2\sigma^2)},\qquad g(x,y;\sigma) = \frac{1}{2\pi\sigma^2}\,e^{-(x^2+y^2)/(2\sigma^2)}. You discretize the Gaussian by sampling e−(n2+m2)/(2σ2)e^{-(n^2+m^2)/(2\sigma^2)} on the integer grid and renormalizing to unit sum. Samples beyond ±3σ\pm 3\sigma are negligible, so a radius of ⌈3σ⌉\lceil 3\sigma\rceil suffices. The Gaussian is the only circularly symmetric kernel that is also separable: g(x,y)=g(x) g(y)g(x,y) = g(x)\,g(y). Filtering with two 1D passes costs O(2N)O(2N) per pixel instead of O(N2)O(N^2).
Output from cell 14 The book shows the same progression on a photograph of a zebra.

Properties of the Gaussian

  1. Its Fourier transform is another Gaussian, G(ω;σ)=e−ω2σ2/2G(\omega;\sigma)=e^{-\omega^2\sigma^2/2}, which is monotonically decreasing, with no side lobes (contrast Figure 17.3b). A wider Gaussian in space is a narrower one in frequency.
  2. Composition adds variances: g(σ1)∗g(σ2)=g(σ3)g(\sigma_1)*g(\sigma_2)=g(\sigma_3) with σ32=σ12+σ22\sigma_3^2=\sigma_1^2+\sigma_2^2. Blurring twice is blurring once by a larger σ\sigma.
Both properties hold exactly only in the continuous case; the sampled kernel satisfies them to a close approximation.
Output from cell 16

Blur as a perceptual low-pass: the block portrait

Harmon and Julesz’s famous demonstration: a face quantized into coarse blocks is hard to read, because the block edges inject high-frequency energy that the visual system latches onto. Low-pass filtering (a Gaussian blur, or simply squinting) removes those spurious high frequencies and the face re-emerges. The book shows the original block portrait of Lincoln. Below you build a block portrait from a photograph by averaging it over 16×1616\times16 blocks, then blur it.
Output from cell 19

Binomial filters

A binomial filter is what you get by convolving the elementary two-tap averager [1,1][1,1] with itself nn times. The coefficients are the nn-th row of Pascal’s triangle: b1=[1,1],b2=[1,2,1],b3=[1,3,3,1],…b_1=[1,1],\quad b_2=[1,2,1],\quad b_3=[1,3,3,1],\quad\dots Key facts (discrete analogues of the Gaussian’s):
  • DC gain =∑bn=2n=\sum b_n = 2^n, so normalize by 2n2^n.
  • Variance σn2=n/4\sigma_n^2 = n/4.
  • Composition: bn∗bm=bn+mb_n * b_m = b_{n+m} and σn2+σm2=σn+m2\sigma_n^2+\sigma_m^2=\sigma_{n+m}^2.
  • The Fourier magnitude B2n(u)=(2+2cos⁡(2πu/N))nB_{2n}(u)=(2+2\cos(2\pi u/N))^n is zero-phase and monotonic: a discrete filter with no side lobes.
Output from cell 21

2D binomial filters and perfect cancellation of the checkerboard

By separability the 2D binomial is the outer product b2,2=[1,2,1]⊤[1,2,1]b_{2,2}=[1,2,1]^\top[1,2,1], i.e. 116[121242121].\tfrac{1}{16}\begin{bmatrix}1&2&1\\2&4&2\\1&2&1\end{bmatrix}. The highest representable frequency is the alternating wave […,1,−1,1,−1,… ][\dots,1,-1,1,-1,\dots]. Convolving it with [1,2,1]/4[1,2,1]/4 gives (1−2+1)/4=0(1-2+1)/4 = 0, so the binomial annihilates the checkerboard exactly. The box [1,1,1]/3[1,1,1]/3 leaves a residual (1−1+1)/3=1/3(1-1+1)/3 = 1/3. Figure 17.8 shows this on an image corrupted by a checkerboard pattern. The book uses a photograph of a boat.
Output from cell 23

Binomials converge to the Gaussian

Repeatedly convolving [1,1][1,1] is repeated averaging, so by the central limit theorem the normalized binomial bn/2nb_n/2^n approaches a Gaussian of variance n/4n/4 as nn grows. This is why small binomials (e.g. [1,2,1][1,2,1], [1,4,6,4,1][1,4,6,4,1]) are the standard cheap, integer-arithmetic Gaussian approximations used in image pyramids.
Output from cell 26

Concluding remarks

Three blur filters, one theme, a low-pass average: The box is fast but its ripples pass spurious high frequencies. The Gaussian is the ideal isotropic low-pass and composes cleanly under σ2\sigma^2 addition. The binomial is the practical, integer-arithmetic bridge, a discrete filter that keeps the Gaussian’s good behavior and, by the central limit theorem, becomes a Gaussian in the limit. These are the building blocks for downsampling, upsampling, and image pyramids in the chapters that follow.