Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #78), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 19 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. A video is a 3D volume ℓ(x,y,t)\ell(x,y,t). Once you think of it that way, motion becomes orientation: a static point is a vertical line in the xx-tt slice, and a point moving at velocity vv is a line of slope 1/v1/v. This section builds the chapter’s space-time tools on a real pedestrian video, a static-camera clip of people crossing a plaza (OpenCV’s vtest.avi sample), the same kind of scene the book uses (Figure 19.1). It covers the xx-tt view of motion, its space-time Fourier signature, the spatiotemporal Gaussian and its velocity-skewed form, velocity-tuned blur, spatiotemporal derivatives, and the velocity-nulling filter that erases objects moving at a chosen velocity.

A video is a space-time volume

Stacking the frames along tt gives a 3-D volume. Slice it at a fixed row mm and you get an xx-tt image: the static background is made of vertical streaks (same xx for every tt), while each walking person traces a diagonal streak whose slope is its velocity. This is the whole idea of the chapter: motion has become orientation.
Output from cell 5 Output from cell 6

Motion is a slanted plane in the Fourier domain

A globally translating image ℓ(x,y,t)=ℓ0(x−vxt, y−vyt)\ell(x,y,t)=\ell_0(x-v_xt,\,y-v_yt) has all its energy on the plane wt+vxwx+vywy=0.w_t + v_x w_x + v_y w_y = 0. In 1-D space, a pulse moving at velocity vv is a slanted band in xx-tt, and its 2-D Fourier transform is a sinc ridge lying along the line wt+v wx=0w_t+v\,w_x=0: vertical for a static pulse, tilting as the speed grows.
Output from cell 8

The spatiotemporal Gaussian

The separable space-time Gaussian g(x,y,t;σ,σt)=1(2π)3/2σ2σte−(x2+y2)/2σ2 e−t2/2σt2g(x,y,t;\sigma,\sigma_t)=\tfrac{1}{(2\pi)^{3/2}\sigma^2\sigma_t} e^{-(x^2+y^2)/2\sigma^2}\,e^{-t^2/2\sigma_t^2} is an isotropic blob in xx-tt. Skewing it along a velocity, g(x−vxt, y−vyt, t)g(x-v_xt,\,y-v_yt,\,t), tilts the blob so its long axis follows that motion. Convolving with the skewed kernel is what blurs along a velocity.
Output from cell 9

Velocity-tuned (temporal) blur

Averaging the volume along a velocity keeps whatever moves at that velocity sharp (it sits still in the motion-compensated stack) while everything else smears. Tuning to v=0v=0 keeps the static background crisp and blurs the walkers; tuning to a walker’s velocity makes that walker snap into focus while the background streaks.
Output from cell 10

Spatiotemporal Gaussian derivatives

The space-time gradient ∇g=(gx,gy,gt)\nabla g=(g_x,g_y,g_t) gives oriented derivative filters. The temporal derivative gt=−tσt2gg_t=-\tfrac{t}{\sigma_t^2}g responds to change over time: it is large exactly where something moves and zero on the static background.
Output from cell 11 Output from cell 11

The velocity-nulling filter

By the brightness-constancy relation, an image moving at exactly (vx,vy)(v_x,v_y) satisfies ∂tℓ+vx∂xℓ+vy∂yℓ=0\partial_t\ell + v_x\partial_x\ell + v_y\partial_y\ell = 0. So the filter h=gt+vxgx+vygyh = g_t + v_x g_x + v_y g_y annihilates anything moving at (vx,vy)(v_x,v_y) while passing everything else. Nulling v=0v=0 removes the static background (only the walkers survive); nulling a walker’s velocity erases that walker while the rest remain.
Output from cell 12 Output from cell 12

Concluding remarks

Treating time as a third axis turns motion problems into geometry: the same Gaussian, derivative, and steering ideas from the spatial chapters carry over, and the brightness-constancy constraint becomes a single linear filter that can select or reject a velocity. These are the foundations for motion estimation and optical flow.