SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
Researchers have introduced SeePhys Pro, a new benchmark designed to evaluate how well artificial intelligence models maintain reasoning capabilities when critical information shifts from text to images in physics problems. The study reveals that current frontier models struggle with representation invariance, showing performance degradation as visual elements increase, with visual variable grounding identified as a major bottleneck. To address this inference-time fragility, the team developed large training corpora for multimodal Reinforcement Learning with Verifiable Rewards (RLVR). They employed blind training, where all training images were masked, as a diagnostic control. Surprisingly, this method still improved performance on unmasked validation sets. Further analysis using text-deletion and image-mask-rate controls suggests these gains stem from residual textual and distributional cues rather than genuine visual evidence. The findings emphasize the necessity of evaluating multimodal reasoning not just by final-answer accuracy, but also by robustness under modality transfer and diagnostics that verify reliance on task-critical visual data.
Wire timeline
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
Researchers have introduced SeePhys Pro, a new benchmark designed to evaluate how well artificial intelligence models maintain reasoning capabilities when critical information shifts from text to images in physics problems. The study reveals that current frontier models struggle with representation invariance, showing performance degradation as visual elements increase, with visual variable grounding identified as a major bottleneck. To address this inference-time fragility, the team developed large training corpora for multimodal Reinforcement Learning with Verifiable Rewards (RLVR). They employed blind training, where all training images were masked, as a diagnostic control. Surprisingly, this method still improved performance on unmasked validation sets. Further analysis using text-deletion and image-mask-rate controls suggests these gains stem from residual textual and distributional cues rather than genuine visual evidence. The findings emphasize the necessity of evaluating multimodal reasoning not just by final-answer accuracy, but also by robustness under modality transfer and diagnostics that verify reliance on task-critical visual data.
cs.AI updates on arXiv.org