ReplaySCM: A New Benchmark for Executable Causal Mechanism Induction
Researchers have introduced ReplaySCM, a novel benchmark designed to evaluate the ability of language models to induce executable causal mechanisms from finite interventional evidence. Unlike existing benchmarks that focus on local answers or graph structures, ReplaySCM consists of 1,300 items based on binary worlds generated by latent acyclic Boolean structural causal models. Systems must output mechanism maps in a restricted Boolean DSL, which are then parsed and replayed on training and held-out intervention worlds to assess behavioral correctness rather than syntactic matching. The benchmark varies structural information disclosure through settings like Ordered, Block-order, and Hidden-roots. Results indicate that while frontier large language models can infer parts of functional-parent structures, their performance drops significantly when order or root structures are hidden. Additionally, the study evaluates a support-audit ladder, demonstrating that increased evidence coverage does not eliminate the gap between ordered and hidden-order scenarios. This work complements current causal reasoning benchmarks by focusing on executable replay generalization without claiming unique identification of the latent model.
Wire timeline
ReplaySCM: A New Benchmark for Executable Causal Mechanism Induction
Researchers have introduced ReplaySCM, a novel benchmark designed to evaluate the ability of language models to induce executable causal mechanisms from finite interventional evidence. Unlike existing benchmarks that focus on local answers or graph structures, ReplaySCM consists of 1,300 items based on binary worlds generated by latent acyclic Boolean structural causal models. Systems must output mechanism maps in a restricted Boolean DSL, which are then parsed and replayed on training and held-out intervention worlds to assess behavioral correctness rather than syntactic matching. The benchmark varies structural information disclosure through settings like Ordered, Block-order, and Hidden-roots. Results indicate that while frontier large language models can infer parts of functional-parent structures, their performance drops significantly when order or root structures are hidden. Additionally, the study evaluates a support-audit ladder, demonstrating that increased evidence coverage does not eliminate the gap between ordered and hidden-order scenarios. This work complements current causal reasoning benchmarks by focusing on executable replay generalization without claiming unique identification of the latent model.
cs.AI updates on arXiv.org