When Can Digital Personas Reliably Approximate Human Survey Findings?
A new academic study published on arXiv investigates the reliability of Large Language Model (LLM)-powered digital personas as substitutes for human survey respondents. Using data from the LISS panel, researchers constructed personas based on background variables and pre-2023 survey histories, testing them against held-out post-cutoff answers. The analysis covered four persona architectures, three LLMs, and two prediction tasks, evaluating performance across question, respondent, distributional, equity, and clustering levels. Results indicate that digital personas improve alignment with human response distributions, particularly in domains involving stable attributes and values. Retrieval-augmented architectures showed the most significant improvements. However, these personas remain limited in predicting individual responses and fail to recover multivariate respondent structures. Performance is heavily dependent on human response structure rather than model choice, excelling in low-variability questions and common patterns while struggling with subjective, heterogeneous, or rare responses. The findings offer practical guidance for researchers on when digital personas are appropriate for survey research and when traditional human validation remains necessary.
Wire timeline
When Can Digital Personas Reliably Approximate Human Survey Findings?
A new academic study published on arXiv investigates the reliability of Large Language Model (LLM)-powered digital personas as substitutes for human survey respondents. Using data from the LISS panel, researchers constructed personas based on background variables and pre-2023 survey histories, testing them against held-out post-cutoff answers. The analysis covered four persona architectures, three LLMs, and two prediction tasks, evaluating performance across question, respondent, distributional, equity, and clustering levels. Results indicate that digital personas improve alignment with human response distributions, particularly in domains involving stable attributes and values. Retrieval-augmented architectures showed the most significant improvements. However, these personas remain limited in predicting individual responses and fail to recover multivariate respondent structures. Performance is heavily dependent on human response structure rather than model choice, excelling in low-variability questions and common patterns while struggling with subjective, heterogeneous, or rare responses. The findings offer practical guidance for researchers on when digital personas are appropriate for survey research and when traditional human validation remains necessary.
cs.AI updates on arXiv.org