On Variance Reduction in Learning Mean Flows
Researchers Juanwu Lu and Ziran Wang from arXiv have published a new study addressing the instability of MeanFlow training in one-step generative modeling. This approach is critical for reducing inference costs in diffusion and flow-matching models. The authors identify that the notorious instability, characterized by non-decreasing loss and unbounded gradient variance, stems from a misuse of the conditional velocity field. Specifically, this field serves dual statistical roles as both an unbiased regression target and a Monte Carlo control variate, but the original loss function assigns an incorrect coefficient to the latter. The paper derives the optimal coefficient in closed form, demonstrating that various concurrent fixes are practical realizations of this same optimum. Experimental results on two-dimensional benchmarks and latent Diffusion Transformers show that using the optimal coefficient improves sample quality by up to 54 percent. Furthermore, the study reveals a mismatch between gradient variance minimization and Fréchet Inception Distance (FID) optimization, indicating that the best FID scores often prefer direct use of conditional velocity despite higher variance.
Wire timeline
On Variance Reduction in Learning Mean Flows
Researchers Juanwu Lu and Ziran Wang from arXiv have published a new study addressing the instability of MeanFlow training in one-step generative modeling. This approach is critical for reducing inference costs in diffusion and flow-matching models. The authors identify that the notorious instability, characterized by non-decreasing loss and unbounded gradient variance, stems from a misuse of the conditional velocity field. Specifically, this field serves dual statistical roles as both an unbiased regression target and a Monte Carlo control variate, but the original loss function assigns an incorrect coefficient to the latter. The paper derives the optimal coefficient in closed form, demonstrating that various concurrent fixes are practical realizations of this same optimum. Experimental results on two-dimensional benchmarks and latent Diffusion Transformers show that using the optimal coefficient improves sample quality by up to 54 percent. Furthermore, the study reveals a mismatch between gradient variance minimization and Fréchet Inception Distance (FID) optimization, indicating that the best FID scores often prefer direct use of conditional velocity despite higher variance.
cs.AI updates on arXiv.org