Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
A new research paper published on arXiv investigates whether embodied vision-language model (VLM) agents possess the cognitive capability for mirror self-recognition, a trait historically associated with higher-order cognition in certain animal species. The authors introduce a controlled 3D benchmark requiring first-person VLM agents to infer hidden body attributes from their reflections while avoiding misattribution to others. To ensure rigorous evaluation, the study tests scenarios involving mirror removal, misleading cues, and occluded reflections, analyzing decision processes through mirror-seeking behavior and reasoning-action consistency. Experimental results indicate that mirror-based self-identification primarily emerges in stronger VLMs, which successfully utilize reflected evidence for action. In contrast, weaker models often inspect mirrors but fail to extract relevant self-information or confuse their reflection with others. The study further demonstrates that self-referential language alone does not constitute grounded self-identification, highlighting conflicts between language and vision. This research provides a critical diagnostic tool for determining if embodied self-grounding in AI is causally rooted in perception and action rather than relying on priors, prompt compliance, or confabulation.
Wire timeline
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
A new research paper published on arXiv investigates whether embodied vision-language model (VLM) agents possess the cognitive capability for mirror self-recognition, a trait historically associated with higher-order cognition in certain animal species. The authors introduce a controlled 3D benchmark requiring first-person VLM agents to infer hidden body attributes from their reflections while avoiding misattribution to others. To ensure rigorous evaluation, the study tests scenarios involving mirror removal, misleading cues, and occluded reflections, analyzing decision processes through mirror-seeking behavior and reasoning-action consistency. Experimental results indicate that mirror-based self-identification primarily emerges in stronger VLMs, which successfully utilize reflected evidence for action. In contrast, weaker models often inspect mirrors but fail to extract relevant self-information or confuse their reflection with others. The study further demonstrates that self-referential language alone does not constitute grounded self-identification, highlighting conflicts between language and vision. This research provides a critical diagnostic tool for determining if embodied self-grounding in AI is causally rooted in perception and action rather than relying on priors, prompt compliance, or confabulation.
cs.AI updates on arXiv.org