Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Researchers have introduced Explicit Visual Premise Verification (EVPV), a novel method designed to enhance the reliability of Vision-Language Process Reward Models (VL-PRMs). Current VL-PRMs often struggle with distinguishing between genuine reasoning errors and perceptual misinterpretations of images, leading to systematic false positives and negatives. EVPV addresses this by decoupling perceptual uncertainty from logical evaluation through a lightweight verification interface. The system prompts policies to generate step-wise visual checklists while independently extracting structured visual constraints from input images. By matching these claims against constraints, EVPV computes a visual reliability signal that calibrates step rewards via reliability gating. Experimental results on VisualProcessBench and six multimodal reasoning benchmarks demonstrate that EVPV significantly improves step-level verification and boosts Best-of-N reranking accuracy compared to strong baselines. Furthermore, causal evidence confirms that performance gains stem from constraint fidelity rather than incidental prompt effects. This advancement offers a robust solution for improving intermediate reasoning scoring in multimodal AI systems without requiring per-step tool calls, marking a significant step forward in vision-language model reliability.
Wire timeline
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Researchers have introduced Explicit Visual Premise Verification (EVPV), a novel method designed to enhance the reliability of Vision-Language Process Reward Models (VL-PRMs). Current VL-PRMs often struggle with distinguishing between genuine reasoning errors and perceptual misinterpretations of images, leading to systematic false positives and negatives. EVPV addresses this by decoupling perceptual uncertainty from logical evaluation through a lightweight verification interface. The system prompts policies to generate step-wise visual checklists while independently extracting structured visual constraints from input images. By matching these claims against constraints, EVPV computes a visual reliability signal that calibrates step rewards via reliability gating. Experimental results on VisualProcessBench and six multimodal reasoning benchmarks demonstrate that EVPV significantly improves step-level verification and boosts Best-of-N reranking accuracy compared to strong baselines. Furthermore, causal evidence confirms that performance gains stem from constraint fidelity rather than incidental prompt effects. This advancement offers a robust solution for improving intermediate reasoning scoring in multimodal AI systems without requiring per-step tool calls, marking a significant step forward in vision-language model reliability.
cs.AI updates on arXiv.org