Flow Matching: From an Average to a Sample

The red image is a prediction; the green image is where the flow ends. Move the orange point to see how they differ.

This text is a temporary draft generated by Codex.

Start with a clean sample \(x_1\) and independent Gaussian noise \(\epsilon\). A training pair traces a straight line between them:

\[x_t=t x_1+(1-t)\epsilon.\]

Several pairs can pass through the same noisy input. With squared-error loss, the model learns their average velocity:

\[v^*(x_t,t)=\mathbb E[x_1-\epsilon\mid x_t,t].\]

The red arrow takes that velocity straight to \(t=1\). Its endpoint is the predicted clean sample:

\[\hat{x}_1=x_t+(1-t)v^*(x_t,t)=\mathbb E[x_1\mid x_t,t].\]

At \(t=0\), noise tells us nothing about the clean sample, so the prediction averages the whole dataset. Later, the possible outcomes narrow. The cloud shows which images are being averaged.

The green sample comes from following the velocity field all the way to the end, updating the direction along the way. At an intermediate time, the red prediction can still be a blend even though the trajectory will end at one image.

The photos illustrate a four-mode Gaussian-mixture flow in 1D. Their weights come from the 1D posterior and are used to average the images. The 500 photos were generated with SDXL-Turbo. Source code.