When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Zihan Chen, Di Zhu, Lei Nico Zheng
Don't use LLM-simulated users for segmentation or targeting decisions. The demographic distortion is structural, not fixable by larger models. If you need synthetic data, benchmark against held-out human responses first.
Teams use LLMs as synthetic users to simulate human survey responses for product and market decisions. No one knows when this substitution is valid versus dangerously wrong.
Method: Four models tested across two survey domains (General Social Survey and World Values Survey) failed in two systematic ways. First, no LLM beat even the strongest non-LLM baseline fit on held-out human data. Second, models over-determined demographics, treating identity as far more predictive of attitudes than it is among real people—present for nearly every question-group combination. On a segment-targeting task, models inflated between-segment gaps two to fourfold, directed teams to the wrong segment in half of U.S. cases and most cross-cultural cases, and manufactured segment splits that don't exist in real people.
Caveats: Tested demographic prompting and standard survey-simulation protocols. Other prompting strategies might perform differently.
Reflections: Can any prompting strategy eliminate the demographic over-determination bias, or is it baked into training data? · Do models trained specifically on survey data (rather than general-purpose LLMs) show the same failures? · At what sample size does real human data become more cost-effective than debugging synthetic-user systems?