Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations
This research paper addresses reliability issues in Visual Grounding models, which often fail when processing semantically mismatched captions by producing approximate bounding boxes that ignore contextual cues. To understand the underlying causes of these counterfactual failures, the authors adopt a mechanistic interpretability approach to examine whether embedding anisotropy is a contributing factor. They introduce a novel similarity-controlled counterfactual caption generation protocol that systematically perturbs object or contextual components within predefined embedding similarity intervals. This method allows for a fine-grained analysis of grounding behavior relative to alignment. Experiments conducted on two Transformer-based models with distinct embedding geometries, specifically BERT-based TransVG and CLIP-based SwimVG, reveal no significant correlation between cosine similarity and approximation errors. The findings indicate that anisotropy alone does not explain these failures, suggesting that future efforts to improve model robustness and faithfulness must investigate finer-grained geometric properties of the embedding space rather than relying solely on current similarity metrics.
Wire timeline
Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations
This research paper addresses reliability issues in Visual Grounding models, which often fail when processing semantically mismatched captions by producing approximate bounding boxes that ignore contextual cues. To understand the underlying causes of these counterfactual failures, the authors adopt a mechanistic interpretability approach to examine whether embedding anisotropy is a contributing factor. They introduce a novel similarity-controlled counterfactual caption generation protocol that systematically perturbs object or contextual components within predefined embedding similarity intervals. This method allows for a fine-grained analysis of grounding behavior relative to alignment. Experiments conducted on two Transformer-based models with distinct embedding geometries, specifically BERT-based TransVG and CLIP-based SwimVG, reveal no significant correlation between cosine similarity and approximation errors. The findings indicate that anisotropy alone does not explain these failures, suggesting that future efforts to improve model robustness and faithfulness must investigate finer-grained geometric properties of the embedding space rather than relying solely on current similarity metrics.
cs.AI updates on arXiv.org