Visual-ERM: Reward Modeling for Visual Equivalence in Vision-to-Code Tasks
Researchers have introduced Visual-ERM, a novel multimodal generative reward model designed to enhance reinforcement learning in vision-to-code tasks. Current Large Vision Language Models often struggle with misaligned reward signals, relying on textual rules or coarse visual embeddings that fail to capture fine-grained discrepancies. Visual-ERM addresses this by providing interpretable, task-agnostic feedback directly in the rendered visual space. When integrated into reinforcement learning, it significantly improved the performance of Qwen3-VL-8B-Instruct, achieving an 8.4-point gain in chart-to-code tasks and consistent improvements in table and SVG parsing. The study also presents VisualCritic-RewardBench, a new benchmark for evaluating image-to-image discrepancies, where the 8B-parameter Visual-ERM outperformed much larger models like Qwen3-VL-235B-Instruct. These findings suggest that fine-grained visual reward supervision is essential for effective vision-to-code reinforcement learning, offering a robust solution to reward hacking and improving test-time scaling through reflection and revision mechanisms.
Wire timeline
Visual-ERM: Reward Modeling for Visual Equivalence in Vision-to-Code Tasks
Researchers have introduced Visual-ERM, a novel multimodal generative reward model designed to enhance reinforcement learning in vision-to-code tasks. Current Large Vision Language Models often struggle with misaligned reward signals, relying on textual rules or coarse visual embeddings that fail to capture fine-grained discrepancies. Visual-ERM addresses this by providing interpretable, task-agnostic feedback directly in the rendered visual space. When integrated into reinforcement learning, it significantly improved the performance of Qwen3-VL-8B-Instruct, achieving an 8.4-point gain in chart-to-code tasks and consistent improvements in table and SVG parsing. The study also presents VisualCritic-RewardBench, a new benchmark for evaluating image-to-image discrepancies, where the 8B-parameter Visual-ERM outperformed much larger models like Qwen3-VL-235B-Instruct. These findings suggest that fine-grained visual reward supervision is essential for effective vision-to-code reinforcement learning, offering a robust solution to reward hacking and improving test-time scaling through reflection and revision mechanisms.
cs.AI updates on arXiv.org