Benchmarking Real-Time Question Answering via Executable Code Workflows
Researchers have introduced RT-QA, a dynamic evaluation framework designed to address the limitations of static benchmarks in assessing real-time information retrieval capabilities of AI agents. Published on arXiv, this study highlights that existing benchmarks fail to capture the temporal dynamics of evolving real-world knowledge. The RT-QA framework utilizes an agent-driven pipeline that autonomously generates executable code for web crawling and DOM-based answer extraction, incorporating a self-repair mechanism to adapt to website structure changes. The benchmark covers 12 domains with 320 Chinese questions across three difficulty levels. Evaluations of state-of-the-art models, including GPT-5.2 and GLM-4.7, revealed significant performance gaps, with top models achieving only 46% accuracy. The analysis identified two primary failure modes: 'Lazy Retrieval,' where agents rely on superficial search snippets, and 'Temporal Confusion,' where agents fail to re-anchor historical data to the current time. These findings underscore the critical need for improved retrieval strategies and robust temporal state management in future search-integrated AI agents.
Wire timeline
Benchmarking Real-Time Question Answering via Executable Code Workflows
Researchers have introduced RT-QA, a dynamic evaluation framework designed to address the limitations of static benchmarks in assessing real-time information retrieval capabilities of AI agents. Published on arXiv, this study highlights that existing benchmarks fail to capture the temporal dynamics of evolving real-world knowledge. The RT-QA framework utilizes an agent-driven pipeline that autonomously generates executable code for web crawling and DOM-based answer extraction, incorporating a self-repair mechanism to adapt to website structure changes. The benchmark covers 12 domains with 320 Chinese questions across three difficulty levels. Evaluations of state-of-the-art models, including GPT-5.2 and GLM-4.7, revealed significant performance gaps, with top models achieving only 46% accuracy. The analysis identified two primary failure modes: 'Lazy Retrieval,' where agents rely on superficial search snippets, and 'Temporal Confusion,' where agents fail to re-anchor historical data to the current time. These findings underscore the critical need for improved retrieval strategies and robust temporal state management in future search-integrated AI agents.
cs.AI updates on arXiv.org