GRIT: Teaching MLLMs to Think with Images
Researchers have introduced GRIT (Grounded Reasoning with Images and Texts), a novel method for training Multimodal Large Language Models (MLLMs) to integrate visual information directly into their reasoning processes. Unlike existing models that rely solely on natural language for reasoning chains, GRIT enables models to interleave text with explicit bounding box coordinates, pointing to specific image regions consulted during analysis. The method utilizes a specialized reinforcement learning approach called GRPO-GR, which rewards final answer accuracy and proper formatting without requiring annotated reasoning chains or bounding box labels. This approach achieves exceptional data efficiency, needing as few as 20 image-question-answer triplets for effective training. Comprehensive evaluations indicate that GRIT successfully unifies reasoning and grounding capabilities, allowing MLLMs to produce coherent, visually grounded reasoning chains. The paper, accepted at NeurIPS 2025, represents a significant advancement in computer vision and artificial intelligence by addressing the limitations of pure natural language reasoning in vision-language tasks.
Wire timeline
GRIT: Teaching MLLMs to Think with Images
Researchers have introduced GRIT (Grounded Reasoning with Images and Texts), a novel method for training Multimodal Large Language Models (MLLMs) to integrate visual information directly into their reasoning processes. Unlike existing models that rely solely on natural language for reasoning chains, GRIT enables models to interleave text with explicit bounding box coordinates, pointing to specific image regions consulted during analysis. The method utilizes a specialized reinforcement learning approach called GRPO-GR, which rewards final answer accuracy and proper formatting without requiring annotated reasoning chains or bounding box labels. This approach achieves exceptional data efficiency, needing as few as 20 image-question-answer triplets for effective training. Comprehensive evaluations indicate that GRIT successfully unifies reasoning and grounding capabilities, allowing MLLMs to produce coherent, visually grounded reasoning chains. The paper, accepted at NeurIPS 2025, represents a significant advancement in computer vision and artificial intelligence by addressing the limitations of pure natural language reasoning in vision-language tasks.
cs.AI updates on arXiv.org