Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving
Researchers propose a novel method to enhance Large Language Models (LLMs) in solving Plane Geometry Problems (PGPS) by decoupling visual understanding from logical reasoning. Addressing the limitation that LLMs cannot directly process visual diagrams, the study introduces a Multimodal LLM (MLLM) Interpreter trained to convert geometric diagrams into concise textual descriptions using Conditional Declaration Language (CDL). This approach avoids end-to-end fine-tuning, which often compromises the base LLM's inherent reasoning capabilities. The MLLM Interpreter is optimized via Chain-of-Thought-augmented Supervised Fine-Tuning and Group Relative Policy Optimization (GRPO), utilizing specific CDL matching rewards for more effective training. The team also constructed a new dataset, Formalgeo7k-Rec-CoT, incorporating manual reviews and CoT annotations. Experimental results on benchmarks like Formalgeo7k-Rec-CoT, Unigeo, and MathVista demonstrate that this method, trained on only 5.5k data points, performs competitively against leading open-source and closed-source MLLMs, highlighting the potential of using concise geometric descriptions as a bridge for multimodal reasoning tasks.
Wire timeline
Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving
Researchers propose a novel method to enhance Large Language Models (LLMs) in solving Plane Geometry Problems (PGPS) by decoupling visual understanding from logical reasoning. Addressing the limitation that LLMs cannot directly process visual diagrams, the study introduces a Multimodal LLM (MLLM) Interpreter trained to convert geometric diagrams into concise textual descriptions using Conditional Declaration Language (CDL). This approach avoids end-to-end fine-tuning, which often compromises the base LLM's inherent reasoning capabilities. The MLLM Interpreter is optimized via Chain-of-Thought-augmented Supervised Fine-Tuning and Group Relative Policy Optimization (GRPO), utilizing specific CDL matching rewards for more effective training. The team also constructed a new dataset, Formalgeo7k-Rec-CoT, incorporating manual reviews and CoT annotations. Experimental results on benchmarks like Formalgeo7k-Rec-CoT, Unigeo, and MathVista demonstrate that this method, trained on only 5.5k data points, performs competitively against leading open-source and closed-source MLLMs, highlighting the potential of using concise geometric descriptions as a bridge for multimodal reasoning tasks.
cs.AI updates on arXiv.org