Large language models can predict the results of social science experiments.
- 8 cites
GPT-4 accurately predicts social science experiment outcomes with correlations comparable to human forecasters, even for studies published after its training data cutoff.
- Why it matters: Understanding whether large language models can reliably forecast experimental results addresses a key gap in applying AI to social science research and practice, potentially improving efficiency and decision-making.
- What they did: The researchers built an archive of 70 preregistered, nationally representative US survey experiments with 469 effects and 119,330 participants, prompting GPT-4 to simulate responses and infer treatment effects.
- The result: GPT-4’s predictions showed strong correlations with actual effects, enabling applications like pilot testing and intervention selection, but also highlighted risks such as effect size overestimation and potential misuse.