Skip to main content
Open In Colab This section was written by Ruimeng Yang (pull request #48), with help from an AI coding agent on the code. It reproduces the ideas of Chapter 46 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. The book’s own figures are not reproduced here, because the book’s license covers only the work in full; links point to them instead. Motion estimation asks a simple question: if something moves between two frames, can you recover where it went? In this section you:
  • build intuition for motion in image coordinates,
  • recover motion with a readable patch-matching baseline,
  • estimate a sparse motion field on synthetic data with known ground truth,
  • measure accuracy with endpoint error and match ratios,
  • study how patch size, search radius, and noise change the results,
  • and inspect failure modes such as repetitive texture and motion that is too large.

The book’s figures

The book introduces the task with two frames of a street in Palma (Figure 46.3): each region of the first frame has to be assigned a displacement into the second. The code in this section solves the same task on synthetic frames with known ground truth. Related figures in the book:

Intuition

Between two video frames, a moving object often keeps a similar local appearance. So a small patch around a point in Frame 1 may reappear a few pixels away in Frame 2. This section uses the image-coordinate convention common in vision:
  • xx increases to the right,
  • yy increases downward,
  • a motion vector is written as (dx,dy)(dx, dy).
So (dx=5,dy=−4)(dx=5, dy=-4) means “move 5 pixels right and 4 pixels up.” The first figure below is generated from the synthetic frames used throughout the section. It marks three landmarks: the source patch in Frame 1, the true target in Frame 2, and the estimated target recovered by local patch matching. Because the example uses a known translation, you can check that the visual displacement and the recovered motion agree exactly.

Synthetic example setup

Start with clean synthetic frames, because the true motion is known exactly. That lets you debug the estimator before dealing with real video. The scene uses a few textured geometric shapes so patch matching has enough visual structure to lock onto. The helper below builds a frame pair with a known integer translation. The motion-intuition figure comes later, once the patch matcher has also produced the estimated target.
The baseline method is deliberately simple:
  1. extract an odd-sized patch around a point in Frame 1,
  2. search a square window around the same location in Frame 2,
  3. score each candidate with SSD (sum of squared differences),
  4. keep the displacement with the lowest score.
It is easy to read and easy to reason about, which makes it a good baseline. The diagnostic figure below shows the whole loop in one place: source patch, searched region, SSD surface, matched patch, and patch error. The next cell defines helpers for patch extraction, SSD scoring, and exhaustive local search.

Single-point motion estimation

First, track one carefully chosen point so every part of the process is visible. Read the top row from left to right: the source patch in Frame 1, the searched region in Frame 2, and the SSD surface over candidate (dx,dy)(dx, dy) displacements. Then compare the bottom-row patches: a low-error match means the winning displacement is also visually plausible.
Output from cell 10
Output from cell 10

Sparse motion field estimation

One point builds intuition, but a motion field needs many arrows. You sample textured points on a grid, run the same local matcher at each one, and compute metrics over all valid points. To keep the figure readable, it shows only a subset of arrows, and exact estimates and mismatches are styled differently so the coherent global translation stands out.
Output from cell 14

Validation metrics

Because the synthetic translation is known, you can score the estimates directly.
  • Endpoint error (EPE) is the Euclidean distance between the estimated vector and the true vector.
  • Exact match ratio is the fraction of estimates that recover the exact integer motion.
  • Correct-within-1-pixel ratio is a softer measure that counts near misses as acceptable.
These metrics answer slightly different questions, so look at all of them. The printed summary reports the mean EPE explicitly.
Output from cell 16

Parameter sweeps

Three knobs matter immediately:
  • Patch size: larger patches are often more stable, but they blur away local detail.
  • Search radius: larger windows can recover larger motion, but they cost more and may invite distractors.
  • Noise level: noisy images reduce the reliability of appearance matching.
The sweeps below come from controlled synthetic experiments. Because the true motion is (5,−4)(5, -4), the search-radius plot also shows when the window first becomes large enough to contain the correct answer. The noise sweep averages over several fixed seeds, so the trend is not an artifact of one noisy draw.
Output from cell 18

Failure cases

A basic local matcher has clear blind spots. Two of the most important:
  1. Repetitive texture: many candidate patches look equally good.
  2. Motion outside the search radius: the correct answer is never even evaluated.
Both panels use the same ingredients as the successful single-point demo. What changes is the geometry of the problem: either the cost surface becomes ambiguous, or the true motion falls outside the searched range. The large-motion panel labels the true target as outside the searched region, so the wrong match is a limit of the search, not a bug in the plot.
Output from cell 20
Output from cell 20

Limitations

Patch matching works well when:
  • motion is moderate,
  • the search window covers the true displacement,
  • local texture is distinctive,
  • and noise is not too strong.
It struggles when texture is ambiguous, motion is too large, illumination changes, or the motion varies strongly inside one patch. These limitations are why more advanced methods exist: Lucas-Kanade, Horn-Schunck optical flow, pyramidal search, and learned flow estimators. As an exercise, replace the global translation with a rotation or a locally varying motion and see where the baseline breaks first.