ProactBench: A New Benchmark for Evaluating Conversational Proactivity in LLMs
Researchers have introduced ProactBench, a novel benchmark designed to evaluate the conversational proactivity of Large Language Models (LLMs). Unlike traditional benchmarks that assess responses to explicit user requests, ProactBench measures an AI's ability to identify and act on implied but unspoken user needs. The framework categorizes this proactivity into three types: Emergent (inference from single anchors), Critical (synthesis across multiple anchors), and Recovery (providing forward-looking value post-task). To ensure robust evaluation, the benchmark employs a multi-agent system comprising a Planner, User Agent, and Assistant Model, utilizing information asymmetries to prevent scoring biases such as style confounding or context leakage. The released corpus features 198 curated dialogues with 624 trigger points across 24 psychometric communication styles, validated by an independent LLM judge. Testing across 16 frontier and open-weight models revealed that the Recovery phase is particularly challenging and poorly predicted by existing standard benchmarks. This study highlights a significant gap in current AI evaluation methods and offers a new signal for assessing advanced conversational capabilities in artificial intelligence systems.
Wire timeline
ProactBench: A New Benchmark for Evaluating Conversational Proactivity in LLMs
Researchers have introduced ProactBench, a novel benchmark designed to evaluate the conversational proactivity of Large Language Models (LLMs). Unlike traditional benchmarks that assess responses to explicit user requests, ProactBench measures an AI's ability to identify and act on implied but unspoken user needs. The framework categorizes this proactivity into three types: Emergent (inference from single anchors), Critical (synthesis across multiple anchors), and Recovery (providing forward-looking value post-task). To ensure robust evaluation, the benchmark employs a multi-agent system comprising a Planner, User Agent, and Assistant Model, utilizing information asymmetries to prevent scoring biases such as style confounding or context leakage. The released corpus features 198 curated dialogues with 624 trigger points across 24 psychometric communication styles, validated by an independent LLM judge. Testing across 16 frontier and open-weight models revealed that the Recovery phase is particularly challenging and poorly predicted by existing standard benchmarks. This study highlights a significant gap in current AI evaluation methods and offers a new signal for assessing advanced conversational capabilities in artificial intelligence systems.
cs.AI updates on arXiv.org