Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
Researchers have introduced Ambig-DS, a new benchmark designed to evaluate how data-science agents handle task-framing ambiguity as they transition from co-pilots to autonomous systems. The study highlights 'silent misframing' as a critical failure mode, where agents produce executable but incorrect results by committing to unintended task interpretations. Ambig-DS comprises two diagnostic suites: Ambig-DS-Target for prediction-target ambiguity and Ambig-DS-Objective for evaluation-objective ambiguity, built upon DSBench and MLE-bench respectively. Testing across five AI agents revealed that ambiguity significantly lowers performance, with failures manifesting as wrong-target submissions rather than execution errors. While allowing agents to ask clarifying questions recovers much of the lost performance, current models struggle to determine when such questions are necessary, often over-asking on clear tasks or remaining silent on ambiguous ones. The findings suggest that recognizing underspecification, rather than just pipeline execution, is the primary bottleneck in current data-science agent evaluations, urging a shift in benchmarking standards to address these nuanced cognitive failures.
Wire timeline
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
Researchers have introduced Ambig-DS, a new benchmark designed to evaluate how data-science agents handle task-framing ambiguity as they transition from co-pilots to autonomous systems. The study highlights 'silent misframing' as a critical failure mode, where agents produce executable but incorrect results by committing to unintended task interpretations. Ambig-DS comprises two diagnostic suites: Ambig-DS-Target for prediction-target ambiguity and Ambig-DS-Objective for evaluation-objective ambiguity, built upon DSBench and MLE-bench respectively. Testing across five AI agents revealed that ambiguity significantly lowers performance, with failures manifesting as wrong-target submissions rather than execution errors. While allowing agents to ask clarifying questions recovers much of the lost performance, current models struggle to determine when such questions are necessary, often over-asking on clear tasks or remaining silent on ambiguous ones. The findings suggest that recognizing underspecification, rather than just pipeline execution, is the primary bottleneck in current data-science agent evaluations, urging a shift in benchmarking standards to address these nuanced cognitive failures.
cs.AI updates on arXiv.org