Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #77), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 18 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. The image derivative is the workhorse of low-level vision: it turns intensity changes into signals you can measure. This section builds the derivative operators of the book’s chapter: the two-tap [1,−1][1,-1] and centered [1,0,−1][1,0,-1] kernels, Gaussian derivatives (through Hermite polynomials), derivative-of-binomial kernels, the Roberts and Sobel operators, the Laplacian and Laplacian-of-Gaussian, unsharp masking, and a working Retinex decomposition. Each comes with a numerical check. The book demonstrates these operators on its own photographs. Here they run on scikit-image test images and on synthetic patterns, with a link to each of the book’s figures. Following the book, derivative maps are shown signed around mid-gray, per channel, so color inputs show colored edges.
Images replaced for licensing. The book is published under a CC BY-NC-ND license, which covers only the book as a whole and not its individual images, so this page does not republish the book’s photographs. License-free images stand in for them:If you use the book for non-commercial purposes, you can swap the originals back in: each line that loads a stand-in carries the original’s link in a comment.

Discretizing the image derivative

The continuous partial derivative ∂ℓ/∂x\partial\ell/\partial x becomes a finite difference. Two choices dominate: d0=[1, −1] (ℓ[n]−ℓ[n−1]),d1=12[1, 0, −1] (ℓ[n+1]−ℓ[n−1]2).d_0 = [1,\,-1]\ (\ell[n]-\ell[n-1]),\qquad d_1 = \tfrac12[1,\,0,\,-1]\ \bigl(\tfrac{\ell[n+1]-\ell[n-1]}{2}\bigr). d0d_0 is a half-pixel-shifted difference; d1d_1 is centered (no shift) and a touch smoother. Applying d1d_1 across xx and down yy gives the two derivative images of Figure 18.2 (the book’s version). Following the book, the derivative of each color channel is shown signed around mid-gray: light where intensity rises and dark where it falls, which paints edges in blue/orange.
Output from cell 7

What the two kernels do in frequency

An ideal derivative multiplies each frequency by jωj\omega, i.e. magnitude grows linearly with frequency. The DFTs D0[u]=1−e−2πju/N, ∣D0∣=2sin⁡(πu/N),D1[u]=jsin⁡(2πu/N),D_0[u] = 1-e^{-2\pi j u/N},\ |D_0| = 2\sin(\pi u/N),\qquad D_1[u] = j\sin(2\pi u/N), both approximate ∣ω∣|\omega| at low frequencies. ∣D0∣|D_0| tracks the ideal further up the band; ∣D1∣|D_1| rolls off earlier (it suppresses the highest frequencies, so it is smoother and less noisy).
Output from cell 10

Gaussian derivatives beat noise

Differentiation amplifies high frequencies, so a raw [1,−1][1,-1] derivative of a noisy image is dominated by noise. Because differentiation and convolution commute, ∂ℓ∂x∗g=ℓ∗∂g∂x,\frac{\partial \ell}{\partial x}*g = \ell * \frac{\partial g}{\partial x}, you can differentiate the smooth Gaussian instead of the noisy image. The first Gaussian derivative gx=−xσ2gg_x=-\tfrac{x}{\sigma^2}g smooths and differentiates in one pass. Figure 18.6 contrasts the two on a noisy photo (the book’s version).
Output from cell 12

Gaussian derivatives and the Hermite family

Higher derivative orders of the Gaussian are Hermite polynomials times the Gaussian: gxn(x;σ)=(−1σ2)nHn ⁣(xσ2) g(x;σ),Hn(x)=2xHn−1−2(n−1)Hn−2.g_{x^n}(x;\sigma) = \Bigl(\tfrac{-1}{\sigma\sqrt2}\Bigr)^n H_n\!\Bigl(\tfrac{x}{\sigma\sqrt2}\Bigr)\,g(x;\sigma),\qquad H_n(x)=2xH_{n-1}-2(n-1)H_{n-2}. Each order adds one more oscillation. Orders 0 to 3 for σ=1\sigma=1:
Output from cell 15

The 2-D Gaussian-derivative triangle

Because the 2-D Gaussian is separable, every mixed partial gxnymg_{x^n y^m} is just the outer product of two 1-D Hermite-weighted Gaussians. Arranged by total order they form a triangle (the top is gg itself; each row adds one derivative). Each kernel is shown signed around mid-gray. These are exactly the oriented center-surround filters a linear front-end computes.
Output from cell 18

Multiscale Gaussian derivatives

The scale σ\sigma selects which edges survive: small σ\sigma picks up fine texture, large σ\sigma only the coarse structure. Here is the Gaussian x-derivative of a photograph at σ=2,4,8\sigma=2,4,8 (the book uses a zebra).
Output from cell 20

Derivatives from binomial filters

Convolving a binomial smoother bnb_n (Pascal’s triangle) with the elementary difference [1,−1][1,-1] gives a family of discrete derivative kernels dn=bn∗[1,−1]d_n=b_n*[1,-1]: smoother as nn grows, all with DC gain 0. d0=[1,−1],d1=[1,0,−1],d2=[1,1,−1,−1],d3=[1,2,0,−2,−1],…d_0=[1,-1],\quad d_1=[1,0,-1],\quad d_2=[1,1,-1,-1],\quad d_3=[1,2,0,-2,-1],\dots
Output from cell 22

Roberts and Sobel in frequency

The Roberts cross (2×22\times2 diagonal differences) and the Sobel-Feldman operator ([1,0,−1][1,0,-1] derivative ×\times [1,2,1][1,2,1] smoothing) are the classic 2-D edge operators. Sobel is separable, Sobelx=[1,0,−1]⊗[1,2,1]⊤\text{Sobel}_x = [1,0,-1]\otimes[1,2,1]^\top, and its smoothing makes it the most isotropic and noise-tolerant. The 2-D DFT magnitudes below show each operator’s directional selectivity.
Output from cell 24

Image gradient and directional derivatives

The gradient ∇ℓ=(∂xℓ,∂yℓ)\nabla\ell=(\partial_x\ell,\partial_y\ell) is a per-pixel vector. The derivative in any direction t=(cos⁡θ,sin⁡θ)\mathbf t=(\cos\theta,\sin\theta) is just a linear combination of the two you already have, with no new convolution needed: ∂ℓ∂t=cos⁡θ ∂xℓ+sin⁡θ ∂yℓ.\frac{\partial\ell}{\partial\mathbf t} = \cos\theta\,\partial_x\ell + \sin\theta\,\partial_y\ell.
Output from cell 26

The Laplacian and Laplacian-of-Gaussian

The Laplacian ∇2ℓ=∂xxℓ+∂yyℓ\nabla^2\ell=\partial_{xx}\ell+\partial_{yy}\ell is the simplest rotationally invariant second-order operator. Smoothed with a Gaussian it becomes the Laplacian-of-Gaussian (the Mexican-hat wavelet), ∇2g=x2+y2−2σ2σ4 g(x,y;σ),\nabla^2 g = \frac{x^2+y^2-2\sigma^2}{\sigma^4}\,g(x,y;\sigma), a center-surround kernel that responds to blobs and zero-crosses at edges. The book shows the second derivatives on a photograph of a wheel; below, a synthetic spoked wheel has edges at every orientation, which is what tests isotropy.
Output from cell 28
Output from cell 30

Sharpening: unsharp masking

Subtracting a blurred copy from twice the image boosts the high frequencies the blur removed: sharpen=2I−b2,2,DC gain=1.\text{sharpen} = 2\mathbf I - b_{2,2},\qquad \text{DC gain} = 1. Applied repeatedly it keeps enhancing edges (until artifacts appear). The photo is in color, so the sharpen kernel is applied to each channel independently. The book uses a photograph of a boat.
Output from cell 33

Retinex: separating reflectance from illumination

An image is reflectance times illumination, ℓ=r⋅l\ell=r\cdot l. In the log domain this is a sum, and Land’s Retinex exploits a statistical gap: reflectance edges are sharp (large log-gradients) while illumination varies smoothly (small gradients). So you threshold the log-gradient: keep the large part as reflectance, integrate it back with a (mirror-padded) Poisson solve, and take the smooth remainder as illumination. The test setup follows the book: a synthetic Mondrian (piecewise-constant reflectance patches) under a smooth, left-bright illumination, laid out as the book’s ℓ(x,y)=r(x,y)⋅l(x,y)\ell(x,y)=r(x,y)\cdot l(x,y) decomposition. Because the ground truth is known, you can measure the recovery: the smooth illumination is recovered almost exactly (correlation ~0.97), and the reflectance comes out flat (~0.77; the residual is faint illumination the single global threshold cannot fully separate).
Output from cell 36

Concluding remarks

From a single idea, differencing neighboring pixels, the chapter builds edge detection, scale selection (Gaussian derivatives), the isotropic Laplacian, sharpening, and gradient-domain reasoning strong enough to separate reflectance from illumination. These operators are the front end of nearly every classical vision pipeline (SIFT, HOG) and echo the center-surround receptive fields of early biological vision.