Mechanistic Study Reveals Hidden States, Not Attention Maps, Predict VLM Reliability
A new study published on arXiv challenges the prevailing assumption that sharp attention maps indicate trustworthiness in Vision-Language Models (VLMs). Researchers analyzed three open-weight VLM families (LLaVA-1.5, PaliGemma, and Qwen2-VL) using a unified mechanistic pipeline called the VLM Reliability Probe. The findings demonstrate that attention structure is a near-zero predictor of correctness, despite being causally necessary for feature extraction. Instead, reliability is more accurately determined by hidden-state geometry and layer-wise margin formation later in the computation process. A single hidden-state linear probe achieved an AUROC greater than 0.95 on the POPE benchmark for two of the three model families. Furthermore, causal neuron-level ablations revealed significant architectural differences: late-fusion models like LLaVA concentrate reliability in fragile late bottlenecks, whereas early-fusion models like PaliGemma and Qwen2-VL distribute reliability widely, showing greater robustness to neuron destruction. These results suggest that monitoring hidden states and sparse late-layer circuits offers a more reliable method for assessing VLM confidence than analyzing attention maps.
Wire timeline
Mechanistic Study Reveals Hidden States, Not Attention Maps, Predict VLM Reliability
A new study published on arXiv challenges the prevailing assumption that sharp attention maps indicate trustworthiness in Vision-Language Models (VLMs). Researchers analyzed three open-weight VLM families (LLaVA-1.5, PaliGemma, and Qwen2-VL) using a unified mechanistic pipeline called the VLM Reliability Probe. The findings demonstrate that attention structure is a near-zero predictor of correctness, despite being causally necessary for feature extraction. Instead, reliability is more accurately determined by hidden-state geometry and layer-wise margin formation later in the computation process. A single hidden-state linear probe achieved an AUROC greater than 0.95 on the POPE benchmark for two of the three model families. Furthermore, causal neuron-level ablations revealed significant architectural differences: late-fusion models like LLaVA concentrate reliability in fragile late bottlenecks, whereas early-fusion models like PaliGemma and Qwen2-VL distribute reliability widely, showing greater robustness to neuron destruction. These results suggest that monitoring hidden states and sparse late-layer circuits offers a more reliable method for assessing VLM confidence than analyzing attention maps.
cs.AI updates on arXiv.org