Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Researchers have introduced Agent-ValueBench, the first comprehensive benchmark dedicated to evaluating the values of autonomous agents, addressing a critical gap in AI safety research. While existing benchmarks focus on Large Language Models (LLMs), this study demonstrates that agent values diverge significantly from their underlying LLMs due to agentic modalities. The benchmark comprises 394 executable environments across 16 domains, featuring 4,335 value-conflict tasks covering 28 value systems and 332 dimensions. Each task was co-synthesized via an end-to-end pipeline and curated by professional psychologists, including golden trajectories for rubric-based judging. By benchmarking 14 frontier models across four mainstream harnesses, the study reveals a "Value Tide" of cross-model homogeneity. Crucially, findings indicate that agent alignment is increasingly influenced by harness configurations and embedded skills rather than just model alignment or prompt steering. This signals a strategic shift in AI safety towards harness alignment and skill steering as primary levers for controlling agent behavior, offering new insights for developing safer, more reliable autonomous systems in diverse operational contexts.
Wire timeline
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Researchers have introduced Agent-ValueBench, the first comprehensive benchmark dedicated to evaluating the values of autonomous agents, addressing a critical gap in AI safety research. While existing benchmarks focus on Large Language Models (LLMs), this study demonstrates that agent values diverge significantly from their underlying LLMs due to agentic modalities. The benchmark comprises 394 executable environments across 16 domains, featuring 4,335 value-conflict tasks covering 28 value systems and 332 dimensions. Each task was co-synthesized via an end-to-end pipeline and curated by professional psychologists, including golden trajectories for rubric-based judging. By benchmarking 14 frontier models across four mainstream harnesses, the study reveals a "Value Tide" of cross-model homogeneity. Crucially, findings indicate that agent alignment is increasingly influenced by harness configurations and embedded skills rather than just model alignment or prompt steering. This signals a strategic shift in AI safety towards harness alignment and skill steering as primary levers for controlling agent behavior, offering new insights for developing safer, more reliable autonomous systems in diverse operational contexts.
cs.AI updates on arXiv.org