Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Researchers have introduced Workspace-Bench 1.0, a new benchmark designed to evaluate the capabilities of AI agents in handling complex workspace tasks involving large-scale file dependencies. Addressing the limitations of existing benchmarks that rely on synthesized or pre-specified files, this study constructs realistic workspaces featuring five worker profiles, 74 file types, and over 20,000 files totaling up to 20GB. The benchmark includes 388 curated tasks, each with specific file dependency graphs, assessed through 7,399 rubrics requiring cross-file retrieval, contextual reasoning, and adaptive decision-making. A lighter version, Workspace-Bench-Lite, is also provided to reduce evaluation costs by approximately 70%. Experimental results involving three popular agent harnesses and five foundation models reveal significant performance gaps; the best-performing agent achieved only about 60% accuracy, falling short of the human baseline of 80.7%, while the average agent performance stood at 45.1%. These findings highlight that current AI agents remain far from reliable in executing complex workspace learning tasks, underscoring the need for further advancements in handling heterogeneous file dependencies and real-world operational contexts.
Wire timeline
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Researchers have introduced Workspace-Bench 1.0, a new benchmark designed to evaluate the capabilities of AI agents in handling complex workspace tasks involving large-scale file dependencies. Addressing the limitations of existing benchmarks that rely on synthesized or pre-specified files, this study constructs realistic workspaces featuring five worker profiles, 74 file types, and over 20,000 files totaling up to 20GB. The benchmark includes 388 curated tasks, each with specific file dependency graphs, assessed through 7,399 rubrics requiring cross-file retrieval, contextual reasoning, and adaptive decision-making. A lighter version, Workspace-Bench-Lite, is also provided to reduce evaluation costs by approximately 70%. Experimental results involving three popular agent harnesses and five foundation models reveal significant performance gaps; the best-performing agent achieved only about 60% accuracy, falling short of the human baseline of 80.7%, while the average agent performance stood at 45.1%. These findings highlight that current AI agents remain far from reliable in executing complex workspace learning tasks, underscoring the need for further advancements in handling heterogeneous file dependencies and real-world operational contexts.
cs.AI updates on arXiv.org