Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
Researchers Guodong Ding and Angela Yao have introduced Gate-and-Merge, a novel zero-shot framework designed for the compositional personalization of vision-language models (VLMs). This approach addresses the challenge of recognizing or describing multiple user-defined concepts jointly at test time without requiring co-occurrence training. The method involves learning each concept independently as a lightweight LoRA adapter paired with a specific token, ensuring the base model remains unchanged and concepts stay disentangled. During inference, the framework merges these concept-specific updates directly in weight space. To prevent interference and suppress irrelevant activations, a gating mechanism estimates textual and visual cues to select only contributing modules. Additionally, the system stabilizes composition by combining only meaningful and mutually consistent updates, thereby preserving individual concept identities. Quantitative and qualitative analyses demonstrate consistent performance improvements across various personalization tasks, including both single-concept and compositional settings. This development represents a significant advancement in artificial intelligence, specifically within computer vision and pattern recognition, offering a more efficient way to customize AI models for complex, multi-concept scenarios.
Wire timeline
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
Researchers Guodong Ding and Angela Yao have introduced Gate-and-Merge, a novel zero-shot framework designed for the compositional personalization of vision-language models (VLMs). This approach addresses the challenge of recognizing or describing multiple user-defined concepts jointly at test time without requiring co-occurrence training. The method involves learning each concept independently as a lightweight LoRA adapter paired with a specific token, ensuring the base model remains unchanged and concepts stay disentangled. During inference, the framework merges these concept-specific updates directly in weight space. To prevent interference and suppress irrelevant activations, a gating mechanism estimates textual and visual cues to select only contributing modules. Additionally, the system stabilizes composition by combining only meaningful and mutually consistent updates, thereby preserving individual concept identities. Quantitative and qualitative analyses demonstrate consistent performance improvements across various personalization tasks, including both single-concept and compositional settings. This development represents a significant advancement in artificial intelligence, specifically within computer vision and pattern recognition, offering a more efficient way to customize AI models for complex, multi-concept scenarios.
cs.AI updates on arXiv.org