Re²Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
Researchers have introduced Re²Math, a new benchmark designed to evaluate the capability of large language models (LLMs) in retrieving mathematical tools from literature during research-level proof construction. While LLMs excel at closed-world reasoning, they often struggle with source-grounded assistance, specifically in identifying existing lemmas and verifying their applicability to specific proof contexts. Re²Math addresses this by creating instances from candidate instrumental citations within main theorem proofs, utilizing hierarchical context and optional anchor hints. The benchmark is citation-agnostic, accepting any admissible theorem that suffices for the proof transition, and employs a release-frozen retrieval artifact to ensure reproducibility. Initial evaluations reveal a significant performance gap; although systems achieve high rates of source grounding, the best fixed-judge Tool Accuracy reaches only 7.0%. This indicates that current AI systems frequently retrieve valid mathematical statements but fail to establish their relevance to local proof steps. By decoupling citation recall, grounding, and proof-gap sufficiency, Re²Math transforms literature-grounded mathematical tool use into a controlled diagnostic task, supporting automatic expansion for continual assessment of AI advancements in mathematical research assistance.
Wire timeline
Re²Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
Researchers have introduced Re²Math, a new benchmark designed to evaluate the capability of large language models (LLMs) in retrieving mathematical tools from literature during research-level proof construction. While LLMs excel at closed-world reasoning, they often struggle with source-grounded assistance, specifically in identifying existing lemmas and verifying their applicability to specific proof contexts. Re²Math addresses this by creating instances from candidate instrumental citations within main theorem proofs, utilizing hierarchical context and optional anchor hints. The benchmark is citation-agnostic, accepting any admissible theorem that suffices for the proof transition, and employs a release-frozen retrieval artifact to ensure reproducibility. Initial evaluations reveal a significant performance gap; although systems achieve high rates of source grounding, the best fixed-judge Tool Accuracy reaches only 7.0%. This indicates that current AI systems frequently retrieve valid mathematical statements but fail to establish their relevance to local proof steps. By decoupling citation recall, grounding, and proof-gap sufficiency, Re²Math transforms literature-grounded mathematical tool use into a controlled diagnostic task, supporting automatic expansion for continual assessment of AI advancements in mathematical research assistance.
cs.AI updates on arXiv.org