Skip to main content
Open In Colab This section was written by Kaushik Kachireddy (pull request #82), with help from an AI coding agent (Claude Code) on the code. It reproduces the ideas of Chapter 30 of Foundations of Computer Vision by Antonio Torralba, Phillip Isola, and William T. Freeman. Representation learning is the forward half of vision: mapping raw pixels to a compact embedding that exposes the underlying factors of a scene. This section builds the book’s demonstrations on its toy colored-shapes dataset:
  1. The dataset (Figure 30.4): 3 shapes by 8 colors, with random size, position, and rotation.
  2. Autoencoders (Figure 30.5): an encoder, a bottleneck, and a decoder. You inspect the embedding through reconstructions, nearest neighbors, and a per-layer probe.
  3. K-means (Figure 30.11): clustering as a discrete representation.
  4. Contrastive learning (Figure 30.14): InfoNCE, where the choice of data augmentation decides whether the embedding encodes color or shape.
The colored-shapes data is procedural in the book too, so it is generated here exactly rather than loaded. Contrastive learning is also the idea behind CLIP, the next section, which applies it across images and text.

The colored-shapes dataset

Each image has one of three shapes (circle, square, triangle) in one of eight colors, with random size, position and rotation. Two independent factors, shape and color, which a representation to disentangle. Output from cell 2

Autoencoders

An autoencoder maps each image through a low-dimensional bottleneck and back, trained to reconstruct its input: min⁡f,g  Ex ∥g(f(x))−x∥2\min_{f,g}\;\mathbb{E}_x\,\lVert g(f(x))-x\rVert^2. The bottleneck f(x)f(x) is forced to keep only what is needed to redraw the image, so it becomes a compact representation. You train a small convolutional autoencoder and then look at what its embedding captures.

Reconstructions and nearest neighbors

The reconstructions confirm the bottleneck preserves shape, color and rough pose. More interesting: images whose embeddings are nearest neighbors tend to share both shape and color: the embedding has organized the data by its true factors.
Output from cell 5

What does each layer encode?

Probe every layer with a 1-nearest-neighbor classifier (train features -> predict a test image’s shape / color). The two factors are read out very differently across depth: shape accuracy rises toward ~99% in the deeper, more semantic features, while color accuracy falls as the network abstracts away low-level appearance: the crossing the book reports (Fig 30.5b). Reproducing it needs the book’s architecture (six conv layers, keeping spatial detail) and a real training budget; a too-shallow model, or downsampling all the way to a single vector, does not show it.
Output from cell 7

K-means clustering

Clustering yields a discrete representation: each point is summarized by the index of its nearest code vector. K-means minimises ∑i∥zai−x(i)∥2\sum_i \lVert z_{a_i} - x^{(i)}\rVert^2 by block-coordinate descent, alternating: (assign) each point to its nearest mean, (update) each mean to the average of its assigned points. The figure shows the first iterations on a 2-D dataset with k=5k=5 (Fig 30.11).
Output from cell 9

Contrastive learning: the augmentation chooses the invariance

Contrastive learning pulls together two augmented views of the same image (a positive pair) while pushing apart different images, via the InfoNCE loss: L=−log⁡ef(x)⊤f(x+)/τef(x)⊤f(x+)/τ+∑ief(x)⊤f(xi−)/τ.\mathcal{L} = -\log \frac{e^{f(x)^\top f(x^+)/\tau}}{e^{f(x)^\top f(x^+)/\tau} + \sum_i e^{f(x)^\top f(x_i^-)/\tau}}. The representation becomes invariant to whatever the augmentation changes. So the choice of augmentation decides which factor survives (embedding dimension M=2M=2, so it can be plotted directly on the unit circle):
  • crop only (TcT_c): keeps color, varies which part of the shape is seen -> the embedding organizes by color;
  • color jitter (TsT_s): destroys color, keeps shape -> the embedding organizes by shape.
Output from cell 11 Output from cell 11

Summary

  • A representation re-expresses pixels in terms of their underlying factors. Here the factors are shape and color, and different methods expose them differently.
  • An autoencoder learns a reusable embedding without labels. Its nearest neighbors respect shape and color, and depth trades low-level appearance (color) for more semantic structure (shape).
  • K-means is the discrete cousin: a representation by cluster index.
  • Contrastive learning makes the control explicit: the augmentation you choose is exactly the invariance you get. Crop-only keeps color; color jitter keeps shape. This is the knob behind modern self-supervised encoders, and behind CLIP, where the two views of a pair are an image and its caption.