Apple researchers have introduced Normalizing Trajectory Models, a generative method designed to produce images with only a few sampling steps while retaining an exact likelihood objective for the full trajectory.
Conventional diffusion models generate output through many small Gaussian denoising updates. Compressing that process into a handful of larger transitions breaks the assumption behind those updates. Existing shortcuts often rely on distillation, consistency training or adversarial objectives, which can improve speed but give up the original likelihood framework.
The new approach models each reverse transition as a conditional normalizing flow, an invertible transformation whose probability can be calculated exactly. Shallow invertible blocks operate within each step, while a deeper parallel predictor coordinates the trajectory. The system can be trained from scratch or initialized from a pretrained flow-matching model.
Its exact trajectory likelihood also supports self-distillation: a lightweight denoiser learns from the score function produced by the model itself. The researchers report that this version matched or outperformed strong text-to-image baselines using four steps. The result is a research benchmark rather than a shipping Apple product, and real-world quality, memory use and speed will depend on implementation and hardware.