Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
A new research paper submitted to arXiv challenges the reliability of Large Language Models (LLMs) used as safety judges for AI agents. The authors argue that current benchmarks incorrectly treat LLM verdicts as ground truth without verifying if decisions depend on agent behavior or merely on the wording of evaluation policies. They introduce 'policy invariance' as a critical reliability property, operationalized through three principles: rubric-semantics invariance, rubric-threshold invariance, and ambiguity-aware calibration. Stress tests on four agent-class judges using ASSEBench and R-Judge datasets reveal a significant failure mode: judges respond equally to meaningful normative shifts and meaningless structural rewrites. Content-preserving policy rewrites flipped up to 9.1% of verdicts, with 18-43% of these flips occurring in unambiguous cases. To address this, the study contributes the Policy Invariance Score and a Judge Card reporting protocol, exposing wide reliability gaps invisible to accuracy-only metrics. The authors release their protocol and code to enable future benchmarks to audit evaluators rigorously rather than trusting them by default, highlighting a crucial step toward robust AI safety assessment.
Wire timeline
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
A new research paper submitted to arXiv challenges the reliability of Large Language Models (LLMs) used as safety judges for AI agents. The authors argue that current benchmarks incorrectly treat LLM verdicts as ground truth without verifying if decisions depend on agent behavior or merely on the wording of evaluation policies. They introduce 'policy invariance' as a critical reliability property, operationalized through three principles: rubric-semantics invariance, rubric-threshold invariance, and ambiguity-aware calibration. Stress tests on four agent-class judges using ASSEBench and R-Judge datasets reveal a significant failure mode: judges respond equally to meaningful normative shifts and meaningless structural rewrites. Content-preserving policy rewrites flipped up to 9.1% of verdicts, with 18-43% of these flips occurring in unambiguous cases. To address this, the study contributes the Policy Invariance Score and a Judge Card reporting protocol, exposing wide reliability gaps invisible to accuracy-only metrics. The authors release their protocol and code to enable future benchmarks to audit evaluators rigorously rather than trusting them by default, highlighting a crucial step toward robust AI safety assessment.
cs.AI updates on arXiv.org