Synthetic Replacements for Human Survey Data? The Perils of Large Language Models

IF 5.5 3区材料科学 Q2 CHEMISTRY, PHYSICAL ACS Applied Energy Materials Pub Date : 2024-05-17 DOI:10.1017/pan.2024.5

James Bisbee, Joshua D. Clinton, C. Dorff, Brenton Kenkel, Jennifer M. Larson

{"title":"Synthetic Replacements for Human Survey Data? The Perils of Large Language Models","authors":"James Bisbee, Joshua D. Clinton, C. Dorff, Brenton Kenkel, Jennifer M. Larson","doi":"10.1017/pan.2024.5","DOIUrl":null,"url":null,"abstract":"\n Large language models (LLMs) offer new research possibilities for social scientists, but their potential as “synthetic data” is still largely unknown. In this paper, we investigate how accurately the popular LLM ChatGPT can recover public opinion, prompting the LLM to adopt different “personas” and then provide feeling thermometer scores for 11 sociopolitical groups. The average scores generated by ChatGPT correspond closely to the averages in our baseline survey, the 2016–2020 American National Election Study (ANES). Nevertheless, sampling by ChatGPT is not reliable for statistical inference: there is less variation in responses than in the real surveys, and regression coefficients often differ significantly from equivalent estimates obtained using ANES data. We also document how the distribution of synthetic responses varies with minor changes in prompt wording, and we show how the same prompt yields significantly different results over a 3-month period. Altogether, our findings raise serious concerns about the quality, reliability, and reproducibility of synthetic survey data generated by LLMs.","PeriodicalId":4,"journal":{"name":"ACS Applied Energy Materials","volume":"62 21","pages":""},"PeriodicalIF":5.5000,"publicationDate":"2024-05-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"4","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ACS Applied Energy Materials","FirstCategoryId":"90","ListUrlMain":"https://doi.org/10.1017/pan.2024.5","RegionNum":3,"RegionCategory":"材料科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"CHEMISTRY, PHYSICAL","Score":null,"Total":0}

引用次数: 4

Abstract

Large language models (LLMs) offer new research possibilities for social scientists, but their potential as “synthetic data” is still largely unknown. In this paper, we investigate how accurately the popular LLM ChatGPT can recover public opinion, prompting the LLM to adopt different “personas” and then provide feeling thermometer scores for 11 sociopolitical groups. The average scores generated by ChatGPT correspond closely to the averages in our baseline survey, the 2016–2020 American National Election Study (ANES). Nevertheless, sampling by ChatGPT is not reliable for statistical inference: there is less variation in responses than in the real surveys, and regression coefficients often differ significantly from equivalent estimates obtained using ANES data. We also document how the distribution of synthetic responses varies with minor changes in prompt wording, and we show how the same prompt yields significantly different results over a 3-month period. Altogether, our findings raise serious concerns about the quality, reliability, and reproducibility of synthetic survey data generated by LLMs.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

人工调查数据的合成替代品？大型语言模型的危险

大型语言模型（LLM）为社会科学家提供了新的研究可能性，但它们作为 "合成数据 "的潜力在很大程度上仍不为人所知。在本文中，我们研究了广受欢迎的大型语言模型 ChatGPT 在恢复民意方面的准确性，促使大型语言模型采用不同的 "角色"，然后为 11 个社会政治团体提供感觉温度计分数。ChatGPT 得出的平均分与我们的基线调查--2016-2020 年美国全国大选研究（ANES）--的平均分非常接近。尽管如此，ChatGPT 的抽样对于统计推断并不可靠：与真实调查相比，回答的变化较小，回归系数往往与使用 ANES 数据获得的等效估计值相差很大。我们还记录了合成回答的分布如何随着提示措辞的细微变化而变化，并展示了同一提示在 3 个月内如何产生显著不同的结果。总之，我们的研究结果引起了人们对由 LLM 生成的合成调查数据的质量、可靠性和可重复性的严重担忧。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

ACS Applied Energy Materials Materials Science-Materials Chemistry

CiteScore

10.30

自引率

6.20%

发文量

1368

期刊介绍： ACS Applied Energy Materials is an interdisciplinary journal publishing original research covering all aspects of materials, engineering, chemistry, physics and biology relevant to energy conversion and storage. The journal is devoted to reports of new and original experimental and theoretical research of an applied nature that integrate knowledge in the areas of materials, engineering, physics, bioscience, and chemistry into important energy applications.

期刊最新文献

Issue Publication Information Issue Editorial Masthead Non-Thermal Pulsed Plasma Synthesis of Carbon Nanomaterials from Hydrocarbons: Morphology and Energy Storage Hydrogen Evolution Reaction Using a Sulfanilamide/Citric Acid Derived N, S-Doped Carbon Dot Solidification Engineering for Thermoelectrics: Figure of Merit and Plasticity