← Back to brief
ResearchOfficialPreprintarXiv Computers and Society

Statistical realism does not guarantee LLMs can estimate treatment effects in social science experiments

A preprint reports that large language models (LLMs) tested on a large-scale, cross-national social science experiment showed only a weak correlation between statistical realism—how closely simulated responses match human data—and the accuracy of estimated treatment effects. In some cases, optimizing for realism actually reduced treatment-effect accuracy, particularly for behavioral outcomes. The findings suggest that using realism as a proxy for treatment-effect accuracy in LLM-generated synthetic data may be unreliable.

Why it matters: This challenges a common assumption in AI-driven social science research and raises concerns about relying on LLM simulations for policy or experimental decisions.

Full story at: arXiv Computers and Society