MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Researchers have introduced MonitoringBench, a new semi-automated red-teaming methodology designed to expose vulnerabilities in coding-agent monitors. The study identifies three major challenges in current security practices: mode collapse in attack generation, a conceive-execute gap in large language models, and the high cost of manual elicitation. To address these, the team developed a pipeline that decomposes attack construction into strategy generation, execution, and refinement. Applied to the BashArena environment, this approach produced a benchmark of 2,644 diverse attack trajectories. Results indicate that current monitoring systems significantly overstate their performance; for instance, the Opus-4.5 monitor's catch rate dropped from 94.9% to 60.3% when tested against these refined attacks. The findings reveal that while frontier monitors detect suspicious actions, they often fail against persuasion tactics or misjudge suspiciousness scores. This work provides both a static benchmark and a reusable methodology for evaluating and improving AI agent security as technology evolves.
Wire timeline
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Researchers have introduced MonitoringBench, a new semi-automated red-teaming methodology designed to expose vulnerabilities in coding-agent monitors. The study identifies three major challenges in current security practices: mode collapse in attack generation, a conceive-execute gap in large language models, and the high cost of manual elicitation. To address these, the team developed a pipeline that decomposes attack construction into strategy generation, execution, and refinement. Applied to the BashArena environment, this approach produced a benchmark of 2,644 diverse attack trajectories. Results indicate that current monitoring systems significantly overstate their performance; for instance, the Opus-4.5 monitor's catch rate dropped from 94.9% to 60.3% when tested against these refined attacks. The findings reveal that while frontier monitors detect suspicious actions, they often fail against persuasion tactics or misjudge suspiciousness scores. This work provides both a static benchmark and a reusable methodology for evaluating and improving AI agent security as technology evolves.
cs.AI updates on arXiv.org