VAE-LFA: Plug-and-Play Low Frequency Alignment to Prevent DiT Editor Drift
Researchers have introduced VAE-LFA, a novel training-free method designed to address semantic drift and quality degradation in multi-turn image editing using Diffusion Transformers (DiTs). While DiTs excel at single-turn edits, repeated modifications often cause progressive misalignment. By analyzing the process in VAE latent space, the team identified that DiTs introduce dominant low-frequency drift, whereas Variational Autoencoders (VAEs) contribute stable reconstruction bias. VAE-LFA mitigates this by decomposing latent discrepancies via low-pass filtering and aligning low-frequency statistics with an exponential moving average of previous rounds. This approach preserves high-frequency details while suppressing accumulated drift. The method is plug-and-play, requiring no retraining, ground-truth priors, or access to diffusion parameters, making it suitable for both white-box and black-box models. For white-box systems, it eliminates redundant VAE round trips; for black-box systems, it uses off-the-shelf VAEs for inter-round alignment. Extensive experiments confirm that VAE-LFA significantly improves semantic consistency and visual fidelity across diverse editing scenarios, including controlled and in-the-wild images, offering a robust solution for advanced AI image manipulation.
Wire timeline
VAE-LFA: Plug-and-Play Low Frequency Alignment to Prevent DiT Editor Drift
Researchers have introduced VAE-LFA, a novel training-free method designed to address semantic drift and quality degradation in multi-turn image editing using Diffusion Transformers (DiTs). While DiTs excel at single-turn edits, repeated modifications often cause progressive misalignment. By analyzing the process in VAE latent space, the team identified that DiTs introduce dominant low-frequency drift, whereas Variational Autoencoders (VAEs) contribute stable reconstruction bias. VAE-LFA mitigates this by decomposing latent discrepancies via low-pass filtering and aligning low-frequency statistics with an exponential moving average of previous rounds. This approach preserves high-frequency details while suppressing accumulated drift. The method is plug-and-play, requiring no retraining, ground-truth priors, or access to diffusion parameters, making it suitable for both white-box and black-box models. For white-box systems, it eliminates redundant VAE round trips; for black-box systems, it uses off-the-shelf VAEs for inter-round alignment. Extensive experiments confirm that VAE-LFA significantly improves semantic consistency and visual fidelity across diverse editing scenarios, including controlled and in-the-wild images, offering a robust solution for advanced AI image manipulation.
cs.AI updates on arXiv.org