Towards Robust Sequential Decomposition for Complex Image Editing
Researchers have introduced a new approach to enhance complex image editing using visual generative models, addressing limitations in current single-turn and sequential editing paradigms. While single-turn editing often fails to parse intricate instructions accurately, sequential editing suffers from compounding errors that degrade image fidelity. To overcome these challenges, the team developed a unified in-context editing framework to balance the benefits of sequential decomposition against error accumulation. They created a synthetic data pipeline to generate editing tasks with varying complexity, curating a large-scale dataset with high-quality decomposed sequences. Fine-tuning models on this synthetic data revealed that properly designed sequential decomposition significantly improves robustness as task complexity increases. Furthermore, the study demonstrates that decomposition skills learned from synthetic tasks can effectively transfer to real-world images through co-training with actual editing data. This sim-to-real generalization offers a promising solution for handling complex, multi-step image editing instructions across broader domains, marking a significant advancement in computer vision and generative AI capabilities.
Wire timeline
Towards Robust Sequential Decomposition for Complex Image Editing
Researchers have introduced a new approach to enhance complex image editing using visual generative models, addressing limitations in current single-turn and sequential editing paradigms. While single-turn editing often fails to parse intricate instructions accurately, sequential editing suffers from compounding errors that degrade image fidelity. To overcome these challenges, the team developed a unified in-context editing framework to balance the benefits of sequential decomposition against error accumulation. They created a synthetic data pipeline to generate editing tasks with varying complexity, curating a large-scale dataset with high-quality decomposed sequences. Fine-tuning models on this synthetic data revealed that properly designed sequential decomposition significantly improves robustness as task complexity increases. Furthermore, the study demonstrates that decomposition skills learned from synthetic tasks can effectively transfer to real-world images through co-training with actual editing data. This sim-to-real generalization offers a promising solution for handling complex, multi-step image editing instructions across broader domains, marking a significant advancement in computer vision and generative AI capabilities.
cs.AI updates on arXiv.org