DRoRAE: Multi-Layer Representation Fusion for Enhanced Visual Tokenization
Researchers have introduced DRoRAE (Depth-Routed Representation AutoEncoder), a novel approach to visual tokenization that addresses the limitations of existing methods which typically extract features only from the last layer of pretrained vision encoders. By discarding intermediate layers, current techniques lose rich hierarchical information and low-level visual details. DRoRAE employs a lightweight fusion module that adaptively aggregates features from all encoder layers using energy-constrained routing and incremental correction. This process creates an enriched latent representation compatible with frozen pretrained decoders. The model utilizes a three-phase decoupled training strategy to optimize fusion and fine-tune the decoder. Experimental results on ImageNet-256 demonstrate significant improvements, reducing reconstruction FID from 0.57 to 0.29 and improving generation FID from 1.74 to 1.65. Additionally, the study identifies a log-linear scaling law between fusion capacity and reconstruction quality, establishing 'representation richness' as a scalable dimension for visual tokenizers similar to vocabulary size in natural language processing. These advancements also show positive transfer effects in text-to-image synthesis tasks.
Wire timeline
DRoRAE: Multi-Layer Representation Fusion for Enhanced Visual Tokenization
Researchers have introduced DRoRAE (Depth-Routed Representation AutoEncoder), a novel approach to visual tokenization that addresses the limitations of existing methods which typically extract features only from the last layer of pretrained vision encoders. By discarding intermediate layers, current techniques lose rich hierarchical information and low-level visual details. DRoRAE employs a lightweight fusion module that adaptively aggregates features from all encoder layers using energy-constrained routing and incremental correction. This process creates an enriched latent representation compatible with frozen pretrained decoders. The model utilizes a three-phase decoupled training strategy to optimize fusion and fine-tune the decoder. Experimental results on ImageNet-256 demonstrate significant improvements, reducing reconstruction FID from 0.57 to 0.29 and improving generation FID from 1.74 to 1.65. Additionally, the study identifies a log-linear scaling law between fusion capacity and reconstruction quality, establishing 'representation richness' as a scalable dimension for visual tokenizers similar to vocabulary size in natural language processing. These advancements also show positive transfer effects in text-to-image synthesis tasks.
cs.AI updates on arXiv.org