Skip to main content

Velocity fields in generative models

The velocity-fields section showed how a 2D wind pattern transports air through space. Generative models use the same mathematical object:
Weather. v(x,y)\mathbf{v}(x, y) tells air where to move. Diffusion and flow matching. v(x,t)\mathbf{v}(x, t) tells probability mass, or a generated sample, where to move.
In weather, the atmosphere determines the wind. Leaves, raindrops, and dust move with the wind field at their current positions. In generative modeling, you design the field. You choose a velocity field that moves particles from Gaussian noise to the data distribution. The particles in flow matching play the role of the leaves. Each particle is a sample carried by the learned field vθv_\theta. It starts as a Gaussian draw and ends as a generated data point. Only the starting positions x0x_0 are Gaussian. Once the field moves them, the particle distribution is no longer Gaussian. At an intermediate time tt, the particles are samples from the deformed distribution ptp_t. At t=1t = 1, they are samples from a distribution that approximates the data. All randomness comes from the initial draw. The ODE then determines each particle’s trajectory. Flow matching learns a continuous-time vector field that transports samples from a simple source distribution to a target data distribution. It generalizes diffusion models. Diffusion fixes a stochastic forward process and learns to reverse it. Flow matching directly parameterizes the deterministic ODE from noise to data. Training uses a regression objective on conditional vector fields. Several generative systems use flow matching for images, video, audio, speech, and molecular structures. Examples include Meta’s Movie Gen, Stable Diffusion 3, and Flux.

Definition

A time-dependent velocity field is a function v:Rd×[0,1]→Rd,(x,t)↦v(x,t)v: \mathbb{R}^d \times [0, 1] \to \mathbb{R}^d, \qquad (x, t) \mapsto v(x, t) that assigns a velocity vector v(x,t)∈Rdv(x, t) \in \mathbb{R}^d to every point xx at every time tt. A particle whose trajectory xt∈Rdx_t \in \mathbb{R}^d is driven by this field obeys the ordinary differential equation dxtdt  =  v(xt,t),x0∼p0(x).\frac{d x_t}{d t} \;=\; v(x_t, t), \qquad x_0 \sim p_0(x). Given an initial sample x0x_0 from a source distribution p0p_0 (typically a standard Gaussian), the ODE produces a unique trajectory {xt}t∈[0,1]\{x_t\}_{t \in [0, 1]}. The map ϕt(x0)  =  xt\phi_t(x_0) \;=\; x_t is called the flow induced by vv.

Distribution transport

The flow transports the entire source density p0p_0 forward in time. At each tt, the pushforward pt  =  (ϕt)♯ p0p_t \;=\; (\phi_t)_\sharp \, p_0 is a probability density on Rd\mathbb{R}^d. It describes the particle distribution at time tt. For a suitable velocity field, the pushforward at t=1t = 1 matches the data distribution: p1  ≈  pdata.p_1 \;\approx\; p_{\text{data}}. The pair (v,pt)(v, p_t) is linked by the continuity equation: ∂pt∂t+∇⋅(pt v)=0.\frac{\partial p_t}{\partial t} + \nabla \cdot \bigl( p_t \, v \bigr) = 0. This equation also describes mass conservation in a fluid. Geometrically, the divergence of the mass flux ptvp_t v determines the local rate of change of density. Designing the velocity field defines a fluid flow that transforms noise into data.

Generation as trajectory integration

Once the learned velocity field vθv_\theta approximates a transport field, you can sample from the model by integrating the ODE numerically. The simplest method is Euler integration with step size Δt\Delta t: xt+Δt  =  xt+Δt⋅vθ(xt,t).x_{t + \Delta t} \;=\; x_t + \Delta t \cdot v_\theta(x_t, t). This is the same loop used for the storm example. At each step, evaluate the field at the current position and move a short distance in that direction. Start from x0∼N(0,I)x_0 \sim \mathcal{N}(0, I) and continue until t=1t = 1. The result is a sample x1x_1 whose distribution approximates pdatap_{\text{data}}. Higher-order solvers (Heun, RK4, adaptive Dormand-Prince) can integrate the same field with fewer steps and lower truncation error. You can choose the solver independently of the velocity field.

Views of a velocity field

A velocity field has three common representations:
  1. Vector field view. At each (x,t)(x, t), draw the arrow v(x,t)v(x, t). Quiver plots are this view.
  2. Streamline view. Fix tt and trace integral curves of v(⋅,t)v(\cdot, t).
  3. Particle view. Release a cloud of particles at t=0t = 0 and follow them as they advect under vv. Trajectory plots are this view.
The same mathematics appears in fluid mechanics, dynamical systems, and optical flow in computer vision. Flow matching applies this geometry to generative modeling. The model learns how probability mass moves through space and time.

The learned velocity field

The network does not denoise or predict a discrete sequence of tokens. It regresses one scalar-valued function per output dimension: vθ(x,t)  ≈  v⋆(x,t)v_\theta(x, t) \;\approx\; v^\star(x, t) where v⋆v^\star is a target velocity field induced by a chosen probability path between p0p_0 and pdatap_{\text{data}}. The flow-matching training objective requires a probability path and its corresponding target velocity. The Lipman et al. references below describe this construction.

References

PyTorch reference