COAST Framework Prevents Visual Aphasia in Vision-Language Models via Adaptive Token Pruning
Researchers have introduced COAST (COntrastive Adaptive Semantic Token Pruning), a novel training-free framework designed to optimize Vision-Language Models (LVLMs). The study challenges existing pruning methods that discard low-attention visual tokens, arguing this causes 'Visual Aphasia,' where models lose visual grounding and rely on language priors. COAST addresses this by treating compression as adaptive semantic routing. It utilizes native cross-modal attention to identify query-specific anchors and estimates contextual dispersion through attention entropy. By adapting the retention trade-off between semantic evidence and spatial context, the framework preserves essential information for compositional reasoning. Experimental results across seven benchmarks demonstrate that COAST reduces visual tokens by 77.8% and achieves a 2.15x latency speedup while retaining 98.64% of original performance. The method consistently outperforms strong pruning baselines and generalizes across multiple LVLM families, offering a robust alternative to one-shot scalar pruning for efficient AI inference.
Wire timeline
COAST Framework Prevents Visual Aphasia in Vision-Language Models via Adaptive Token Pruning
Researchers have introduced COAST (COntrastive Adaptive Semantic Token Pruning), a novel training-free framework designed to optimize Vision-Language Models (LVLMs). The study challenges existing pruning methods that discard low-attention visual tokens, arguing this causes 'Visual Aphasia,' where models lose visual grounding and rely on language priors. COAST addresses this by treating compression as adaptive semantic routing. It utilizes native cross-modal attention to identify query-specific anchors and estimates contextual dispersion through attention entropy. By adapting the retention trade-off between semantic evidence and spatial context, the framework preserves essential information for compositional reasoning. Experimental results across seven benchmarks demonstrate that COAST reduces visual tokens by 77.8% and achieves a 2.15x latency speedup while retaining 98.64% of original performance. The method consistently outperforms strong pruning baselines and generalizes across multiple LVLM families, offering a robust alternative to one-shot scalar pruning for efficient AI inference.
cs.AI updates on arXiv.org