TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning
Researchers have introduced TIDE-Bench, a new holistic and efficient benchmark designed to evaluate Tool-Integrated Reasoning (TIR) methods in large language models. Addressing the lack of high-quality, unified evaluation standards in the field, TIDE-Bench offers three primary advantages. First, it features diverse task settings, combining traditional mathematical reasoning and knowledge-intensive QA with newly designed tasks for tool-grounded experimental design and dynamic interaction. This setup probes models' abilities in complex tool invocation and multi-tool coordination. Second, the benchmark employs a comprehensive, task-aware evaluation protocol that measures final answer quality, process reliability, tool-use efficiency, and inference cost. Third, it constructs high-quality evaluation sets by filtering out low-discrimination instances from existing datasets, thereby reducing costs while focusing on challenging samples. Extensive experiments across multiple foundation models reveal persistent bottlenecks in tool grounding, providing critical insights for future TIR research and development. This work aims to standardize and improve the assessment of LLMs enhanced with external computation and execution capabilities.
Wire timeline
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning
Researchers have introduced TIDE-Bench, a new holistic and efficient benchmark designed to evaluate Tool-Integrated Reasoning (TIR) methods in large language models. Addressing the lack of high-quality, unified evaluation standards in the field, TIDE-Bench offers three primary advantages. First, it features diverse task settings, combining traditional mathematical reasoning and knowledge-intensive QA with newly designed tasks for tool-grounded experimental design and dynamic interaction. This setup probes models' abilities in complex tool invocation and multi-tool coordination. Second, the benchmark employs a comprehensive, task-aware evaluation protocol that measures final answer quality, process reliability, tool-use efficiency, and inference cost. Third, it constructs high-quality evaluation sets by filtering out low-discrimination instances from existing datasets, thereby reducing costs while focusing on challenging samples. Extensive experiments across multiple foundation models reveal persistent bottlenecks in tool grounding, providing critical insights for future TIR research and development. This work aims to standardize and improve the assessment of LLMs enhanced with external computation and execution capabilities.
cs.AI updates on arXiv.org