Training-Free Acceleration of Identity-Preserved Generation via Distilled Backbones
A new research paper published on arXiv introduces a training-free method to significantly accelerate identity-preserved image generation. Traditionally, this process relies on computationally expensive, many-step diffusion backbones. The study demonstrates that for identity-conditioned FLUX generation, a frozen InfuseNet identity adapter trained with the dev model can transfer directly to the distilled schnell backbone without retraining. By simply changing the backbone path and disabling classifier-free guidance, the method reduces latency by 5.9 times compared to the standard 28-step dev baseline. Furthermore, it improves performance metrics, increasing ArcFace identity similarity by 0.028 and decreasing LPIPS by 0.016. Analysis of the denoising trajectory reveals that identity fidelity is established within the first 4-8 steps, while subsequent steps mainly refine visual details. Preliminary tests on SDXL and SD1.5 models show similar diminishing returns after intermediate steps. This approach offers a simple, efficient strategy to enhance the trade-off between efficiency and fidelity in personalized image generation, eliminating the need for costly retraining processes.
Wire timeline
Training-Free Acceleration of Identity-Preserved Generation via Distilled Backbones
A new research paper published on arXiv introduces a training-free method to significantly accelerate identity-preserved image generation. Traditionally, this process relies on computationally expensive, many-step diffusion backbones. The study demonstrates that for identity-conditioned FLUX generation, a frozen InfuseNet identity adapter trained with the dev model can transfer directly to the distilled schnell backbone without retraining. By simply changing the backbone path and disabling classifier-free guidance, the method reduces latency by 5.9 times compared to the standard 28-step dev baseline. Furthermore, it improves performance metrics, increasing ArcFace identity similarity by 0.028 and decreasing LPIPS by 0.016. Analysis of the denoising trajectory reveals that identity fidelity is established within the first 4-8 steps, while subsequent steps mainly refine visual details. Preliminary tests on SDXL and SD1.5 models show similar diminishing returns after intermediate steps. This approach offers a simple, efficient strategy to enhance the trade-off between efficiency and fidelity in personalized image generation, eliminating the need for costly retraining processes.
cs.AI updates on arXiv.org