CIVeX: Causal Intervention Verification for Language Agents
Researchers have introduced CIVeX, a novel causal intervention verifier designed to enhance the reliability of tool-using language agents. While current safeguards like schema validators ensure action validity, they fail to guarantee that state-changing actions produce identifiable causal effects, often leading to reduced utility in confounded workflows. CIVeX addresses this by mapping proposed actions to structural causal queries over a committed action-state graph and checking for identifiability. It issues one of four auditable verdicts: EXECUTE, REJECT, EXPERIMENT, or ABSTAIN, requiring an assumption-scoped causal certificate for execution. Empirical evaluations on Causal-ToolBench demonstrate that CIVeX achieves zero observed false executions under moderate and adversarial confounding, outperforming chain-of-thought LLM verifiers like Claude Opus in utility retention. Additionally, tests on real production logs from IHDP and ZOZO Open Bandit show CIVeX matches oracle correct-execution rates while significantly reducing false executions compared to naive baselines. The study concludes that intervention identifiability, rather than mere action validity, is the critical missing component for dependable autonomous agent tool use.
Wire timeline
CIVeX: Causal Intervention Verification for Language Agents
Researchers have introduced CIVeX, a novel causal intervention verifier designed to enhance the reliability of tool-using language agents. While current safeguards like schema validators ensure action validity, they fail to guarantee that state-changing actions produce identifiable causal effects, often leading to reduced utility in confounded workflows. CIVeX addresses this by mapping proposed actions to structural causal queries over a committed action-state graph and checking for identifiability. It issues one of four auditable verdicts: EXECUTE, REJECT, EXPERIMENT, or ABSTAIN, requiring an assumption-scoped causal certificate for execution. Empirical evaluations on Causal-ToolBench demonstrate that CIVeX achieves zero observed false executions under moderate and adversarial confounding, outperforming chain-of-thought LLM verifiers like Claude Opus in utility retention. Additionally, tests on real production logs from IHDP and ZOZO Open Bandit show CIVeX matches oracle correct-execution rates while significantly reducing false executions compared to naive baselines. The study concludes that intervention identifiability, rather than mere action validity, is the critical missing component for dependable autonomous agent tool use.
cs.AI updates on arXiv.org