LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
Researchers have introduced LAGO, a novel framework designed to enhance zero-shot visual-text alignment in computer vision. Traditional methods often struggle with fine-grained recognition because they rely on whole-image analysis or inefficient, redundant image crops, which increases computational costs and introduces noise. Furthermore, early semantic guidance can lead to a 'prediction loop,' where initial errors bias subsequent localization. LAGO addresses these challenges by first performing class-agnostic object-centric candidate discovery to establish a stable visual foundation. It then applies adaptive language-guided refinement, dynamically controlling semantic guidance strength based on intermediate confidence levels. The framework also employs an object-context dual-channel aggregation strategy to combine object-level, contextual, and full-image evidence. Extensive experiments demonstrate that LAGO achieves state-of-the-art performance on standard zero-shot benchmarks and distribution-shift settings. Crucially, it accomplishes this while requiring significantly fewer candidate regions during inference, offering a more efficient and robust solution for localized visual-text alignment without task-specific supervision.
Wire timeline
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
Researchers have introduced LAGO, a novel framework designed to enhance zero-shot visual-text alignment in computer vision. Traditional methods often struggle with fine-grained recognition because they rely on whole-image analysis or inefficient, redundant image crops, which increases computational costs and introduces noise. Furthermore, early semantic guidance can lead to a 'prediction loop,' where initial errors bias subsequent localization. LAGO addresses these challenges by first performing class-agnostic object-centric candidate discovery to establish a stable visual foundation. It then applies adaptive language-guided refinement, dynamically controlling semantic guidance strength based on intermediate confidence levels. The framework also employs an object-context dual-channel aggregation strategy to combine object-level, contextual, and full-image evidence. Extensive experiments demonstrate that LAGO achieves state-of-the-art performance on standard zero-shot benchmarks and distribution-shift settings. Crucially, it accomplishes this while requiring significantly fewer candidate regions during inference, offering a more efficient and robust solution for localized visual-text alignment without task-specific supervision.
cs.AI updates on arXiv.org