HyperTransport: Amortized Conditioning of T2I Generative Models
Researchers have introduced HyperTransport, a novel hypernetwork framework designed to enhance the control and efficiency of Text-to-Image (T2I) generative models. Addressing the limitations of fragile prompting and costly per-concept fine-tuning, HyperTransport maps embeddings from pretrained encoders like CLIP directly to intervention parameters using an optimal transport loss. This approach amortizes optimization costs, enabling the generation of new interventions in a single forward pass, which is 3,600 to 7,000 times faster than existing methods. The framework supports open-ended concept sets, continuous strength control, and cross-modal conditioning, allowing reference images to steer text-based generation. Validated on DMD2 and Nitro-1-PixArt models across 167 held-out concepts, HyperTransport matches the performance of strong baselines on unseen concepts. Evaluations involving CLIP-based metrics, VLM-as-a-judge assessments, and user studies indicate that both human and automated judges prefer HyperTransport over traditional prompting approximately twice as often, marking a significant advancement in stable and predictable model behavior control.
Wire timeline
HyperTransport: Amortized Conditioning of T2I Generative Models
Researchers have introduced HyperTransport, a novel hypernetwork framework designed to enhance the control and efficiency of Text-to-Image (T2I) generative models. Addressing the limitations of fragile prompting and costly per-concept fine-tuning, HyperTransport maps embeddings from pretrained encoders like CLIP directly to intervention parameters using an optimal transport loss. This approach amortizes optimization costs, enabling the generation of new interventions in a single forward pass, which is 3,600 to 7,000 times faster than existing methods. The framework supports open-ended concept sets, continuous strength control, and cross-modal conditioning, allowing reference images to steer text-based generation. Validated on DMD2 and Nitro-1-PixArt models across 167 held-out concepts, HyperTransport matches the performance of strong baselines on unseen concepts. Evaluations involving CLIP-based metrics, VLM-as-a-judge assessments, and user studies indicate that both human and automated judges prefer HyperTransport over traditional prompting approximately twice as often, marking a significant advancement in stable and predictable model behavior control.
cs.AI updates on arXiv.org