When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
A new research paper published on arXiv investigates the root causes of hallucinations in decoder-based Vision-Language Models (VLMs), which are increasingly used in high-stakes applications like medical imaging and autonomous systems. The authors identify a phenomenon called 'geometric over-alignment,' where visual embeddings are excessively aligned with the text manifold to bridge modality gaps. This process injects statistical linguistic bias that overshadows fine-grained visual evidence, leading models to confidently describe non-existent content. Unlike prior works that rely on expensive black-box decoding strategies, this study provides the first quantitative characterization of this issue, showing that linguistic bias concentrates in the top principal components of a universal text subspace. To address this, the researchers propose two complementary remedies: a training-free inference strategy and a bias-aware fine-tuning paradigm. Both methods explicitly project out the biased subspace from visual representations. Experimental results demonstrate significant reductions in hallucinations across POPE, CHAIR, and AMBER benchmarks, along with improved CLAIR scores for long-form captioning. Notably, the training-free variant achieves these improvements without adding computational overhead to the base model.
Wire timeline
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
A new research paper published on arXiv investigates the root causes of hallucinations in decoder-based Vision-Language Models (VLMs), which are increasingly used in high-stakes applications like medical imaging and autonomous systems. The authors identify a phenomenon called 'geometric over-alignment,' where visual embeddings are excessively aligned with the text manifold to bridge modality gaps. This process injects statistical linguistic bias that overshadows fine-grained visual evidence, leading models to confidently describe non-existent content. Unlike prior works that rely on expensive black-box decoding strategies, this study provides the first quantitative characterization of this issue, showing that linguistic bias concentrates in the top principal components of a universal text subspace. To address this, the researchers propose two complementary remedies: a training-free inference strategy and a bias-aware fine-tuning paradigm. Both methods explicitly project out the biased subspace from visual representations. Experimental results demonstrate significant reductions in hallucinations across POPE, CHAIR, and AMBER benchmarks, along with improved CLAIR scores for long-form captioning. Notably, the training-free variant achieves these improvements without adding computational overhead to the base model.
cs.AI updates on arXiv.org