Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
This academic paper addresses a critical limitation in using generative synthetic data for causal inference. While synthetic tabular data is often evaluated based on distributional similarity or predictive performance, these metrics do not guarantee validity for estimating causal effects. The authors demonstrate that fully generative models, including GANs and LLMs, can maintain predictive utility while significantly distorting Average Treatment Effect (ATE) estimates. This structural failure occurs because prediction loss penalizes treatment-effect errors only through overlap-weighted terms, ignoring the need for accurate treatment-effect contrasts. To resolve this, the study proposes a hybrid synthetic-data framework that generates covariates separately from treatment and outcome mechanisms, enabling causal-purpose treatment assignments. Evaluations across synthetic datasets and ACTG experiments show that this hybrid approach improves causal fidelity compared to fully generative baselines. Notably, LLM-based hybrid synthesis outperforms CTGAN in preserving ATE and benchmarking finite-sample estimators. The research offers a robust methodology for researchers relying on synthetic data for causal analysis, highlighting specific remedies for existing pitfalls.
Wire timeline
Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
This academic paper addresses a critical limitation in using generative synthetic data for causal inference. While synthetic tabular data is often evaluated based on distributional similarity or predictive performance, these metrics do not guarantee validity for estimating causal effects. The authors demonstrate that fully generative models, including GANs and LLMs, can maintain predictive utility while significantly distorting Average Treatment Effect (ATE) estimates. This structural failure occurs because prediction loss penalizes treatment-effect errors only through overlap-weighted terms, ignoring the need for accurate treatment-effect contrasts. To resolve this, the study proposes a hybrid synthetic-data framework that generates covariates separately from treatment and outcome mechanisms, enabling causal-purpose treatment assignments. Evaluations across synthetic datasets and ACTG experiments show that this hybrid approach improves causal fidelity compared to fully generative baselines. Notably, LLM-based hybrid synthesis outperforms CTGAN in preserving ATE and benchmarking finite-sample estimators. The research offers a robust methodology for researchers relying on synthetic data for causal analysis, highlighting specific remedies for existing pitfalls.
cs.AI updates on arXiv.org