人工智能代理尚未准备好取代人类在行为研究中的角色

内容来源:https://www.sciencenews.org/article/ai-agents-replace-humans-research
内容总结:
AI“数字孪生”能否取代人类受试者?新研究:效果仍不理想,存在“哈哈镜效应”
近年来,用人工智能模拟人类行为、甚至创建“数字孪生”来替代真人参与社会科学研究,成为部分学者心中的理想方案。真人受试者成本高昂、容易疲劳,也可能因某些研究遭受心理压力。然而,9月2日发表于《科学进展》(Science Advances)的一项新研究表明,这一愿景的实现恐怕还需时日。
研究人员发现,旨在模仿特定个体行为的AI“数字孪生”,实际上会扭曲其所代表个体的真实观点,产生类似“哈哈镜”的失真效应。论文作者之一、纽约哥伦比亚商学院计算社会科学家奥利维尔·图比亚(Olivier Toubia)表示:“这项技术有一定潜力,但(数字孪生的)整体表现令人有些失望。”
在这项发表于《营销科学》(Marketing Science)的早期研究基础上,图比亚团队曾招募了2000多名来自美国各地的受访者,就年龄、族裔、收入、教育、宗教信仰、政治倾向、性格特质、消费习惯、数学能力及词汇量等500多个问题进行了详细问卷调查。受访者还完成了多项用于评估思维模式和认知偏差的线上测试。图比亚表示,由于当前关于数字孪生效用的研究仍有限,该项目旨在建立一个可供学界广泛使用的开源数据集,目前该数据集已被下载约2.5万次。
在本次新研究中,研究人员将每位受访者的全部信息输入大型语言模型,并指示模型模拟该个体进行回答。随后,团队开展了19项社会科学实验,涵盖从人们对两党候选人捐款者的态度、到对算法招聘的看法等多种情境。
结果显示,这些数字孪生的表现虽优于随机水平,但平均约有四分之一的回答是错误的,整体准确率仅与仅获得人口统计信息的聊天机器人相当。不过,与只用人口统计信息的模型相比,数字孪生在捕捉个体间真实差异方面有所进步。例如,在对自我控制能力评分时,真人受访者中一人可能自评为2分,另一人可能是4分;而仅掌握人口统计信息的AI可能给两人都打3分,抹平差异;数字孪生则可能给出3分和5分,虽然仍有偏差,但至少反映群体中的潜在差异。
团队分析认为,数字孪生整体表现不佳主要源于几类系统性失真:其回答趋于同质化,且容易偏向人口统计学的刻板印象;对于收入和教育水平更高的受访者,其模拟准确度会提高;同时,数字孪生表现出某些固定偏差,如对他人的信任度更高、对技术威胁的担忧更少,且整体上比真实人类显得更为“理性”。
宾夕法尼亚州立大学AI研究员、经济学家哈迪·侯赛尼(Hadi Hosseini)也在其关于AI医疗决策的研究中观察到类似现象:“大型语言模型会以特定方向扭曲人类判断,使之变得过于理性、过于合理,与真实的人类行为并不一致。”
不过,侯赛尼认为数字孪生仍有改进空间。他建议,当前研究采用了“非常静态的问题设定”,若能引入更动态的方法,例如让聊天机器人全天候“旁听”个体行为,或与其定期交流,将有望获得更优的数据集。图比亚也计划尝试更复杂的训练方法。但他同时指出,即便是当前这种数字孪生在某些场景下仍有实用价值,例如,当研究人员需要获取详细回答而真人受试者容易疲惫敷衍时,不知疲倦的AI可以生成完整长文;又如,在正式实验前用数字孪生进行预测试,验证实验设计是否合理,从而节省真人受试者有限的精力。
但图比亚也提醒学界保持谦逊:“社会科学通常认为量表和问卷调查能够涵盖人类经验的全貌,但用机器预测人类行为极为困难。我们对合成数据应抱有现实的期望。”
中文翻译:
AI智能体尚未准备好替代人类参与行为研究
一项新研究表明,数字孪生对问题的反应与其所模拟的真实个体并不一致
用AI替代品(即数字孪生)取代人类受试者,正被一些社会科学家列入愿望清单。人类受试者成本高昂、容易疲劳,而且某些研究可能会给他们带来心理困扰。
但一项发表在9月2日《科学进展》杂志上的研究表明,这些愿望可能还需要更长时间才能实现。研究人员指出,旨在模仿特定个体行为的AI孪生,反而似乎扭曲了其对应人类受试者的观点,产生了一种“哈哈镜”效应。
“这种技术有一定的潜力,”纽约哥伦比亚商学院的计算社会科学家奥利维尔·图比亚表示,“但(孪生体的表现)总体而言有些令人失望。”
在去年发表于《营销科学》杂志的研究中,图比亚及其同事招募了来自美国各地的2000多人。这些受访者回答了500多个问题,涉及年龄、种族、收入、教育、宗教习惯、政治偏好、人格特质、消费习惯、数学能力和词汇技能等特征。许多问题来自心理学、经济学和商业研究中常用的量表。受访者还完成了各种旨在评估思维模式和认知偏误的在线测试。
图比亚表示,由于验证数字孪生价值的研究仍然有限,该项目的目的在于创建一套开源数据集供他人使用。“我觉得目前这个数据集已经被下载了大约25000次。”
在最新发表的研究中,为了开发数字孪生,研究人员将每位个体的全部信息输入一个大语言模型,并提示该LLM以那个人的身份进行回答。在19项社会科学实验中,团队评估了从个体及其孪生体如何回应向共和党和民主党候选人捐款的人,到它们对算法招聘“怎么看”等各类问题。
团队发现,这些数字孪生的表现优于随机猜测,但平均约有四分之一的时间会出错。它们与仅接收人口统计信息的聊天机器人的表现大致相当。
不过,与仅了解人口统计信息的LLM相比,数字孪生的确能更好地捕捉个体之间真实反应的差异。例如,一个人可能将自己的自制力评为2分(满分制中),而另一个人评为4分。信息有限的LLM可能会给每个人都打出3分,从而抹平了差异。但数字孪生可能给出3分和5分的结果。虽然仍然不准确,但这些预测能更好地反映群体内潜在差异的分布。
图比亚及其同事将数字孪生的整体不佳表现归因于几个关键的扭曲因素:孪生体的回答往往比人类的回答更加同质化,并且常常偏向人口统计上的刻板印象。同时,当参与者更富有、受教育程度更高时,准确率会上升。孪生体还表现出某些偏误,比如对他人表现出更多的信任,对技术威胁表现出较少的担忧。与它们对应的真人相比,孪生体还显得更加理性。
宾夕法尼亚州立大学的AI研究员兼经济学家哈迪·侯赛尼在自己关于AI智能体在资源稀缺时如何做医疗决策的研究中也看到了类似的现象:“LLM以某种非常特定的方向扭曲了人类判断,朝着非常理性、比真实人类的选择更合乎情理的方向偏移。”
但他表示,数字孪生仍有改进的空间。他指出,研究人员使用了“一套非常静态的问题”。他说,如果加入更多动态方法,例如让聊天机器人在一天中跟随个体或定期与人交谈——这种做法已经相当普遍——可能会产生更好的数据集。
图比亚想尝试更复杂的训练方法。不过他表示,即使本研究中测试的这类数字孪生在某些场景下也是有用的。例如,有时研究人员需要对某个问题的详细回答。疲惫的人类可能只给出单句答案,而不知疲倦的数字孪生可以生成大段文字。同样,在正式实验之前用数字孪生做预测试,可以帮助研究人员确保实验设计有效,然后再消耗人类受试者有限的时间和精力。
但图比亚也呼吁保持谦逊。他说,社会科学家倾向于认为量表和问卷调查能捕捉人类经验的全部范围。但用机器来预测人类行为极其困难。“对于合成数据,我们需要对期望值保持现实态度。”
英文来源:
AI agents aren’t ready to replace humans in behavioral research
A new study shows that digital twins don’t respond like the individuals they were modeled after
Replacing human subjects with AI surrogates, or digital twins, is on some social scientists’ wish lists. Human subjects are expensive, tire easily and can suffer psychological distress from some studies.
But those wishes may take more time to be realized, a study appearing September 2 in Science Advances suggests. AI twins designed to mimic a given individual’s behavior instead seem to distort their surrogate’s views, creating a “funhouse mirror” effect, the researchers note.
“There’s some promise,” says Olivier Toubia, a computational social scientist at Columbia Business School in New York City. “But [the twins’ performance] was overall a bit disappointing.”
In work reported last year in Marketing Science, Toubia and colleagues recruited more than 2,000 individuals from across the United States. Those respondents answered 500-plus questions about characteristics including age, ethnicity, income, education, religious practices, political preferences, personality traits, spending habits, mathematical abilities and vocabulary skills. Many of the questions come from scales that are used commonly in psychology, economic and business research. Respondents also completed various online tests designed to assess thought patterns and biases.
Because research testing the value of digital twins remains limited, the goal of that project was to create an open-source dataset for others to use, Toubia says. “I think it’s been downloaded like 25,000 times at this point.”
To develop the twins in the newly published study, the researchers fed all the information for each individual into a large language model and prompted the LLM to respond as if it were that person. Across 19 social science experiments, the team evaluated everything from how individuals and their twins respond to people who donate to both Republican and Democratic party candidates to what they “think” about algorithmic hiring.
These digital twins performed better than chance, the team found, but they were wrong on average about a quarter of the time. They performed roughly on par with chatbots that received demographic information alone.
The digital twins did, however, better capture real variation in people’s responses compared with the LLMs that knew only demographic info. For example, one person may rate themselves as a 2 on a scale of self-control while another rates themselves as a 4. The LLM with more limited info might report a 3 for each person, washing out any differences. But the digital twin might report a 3 and 5. Though still wrong, the predictions give a better sense of potential differences across the group.
Toubia and colleagues attribute the digital twins’ poor performance overall to several key distortions: Twins’ responses tended to be more homogenous than people’s responses and often skewed to demographic stereotypes. And their accuracy increased with more affluent and educated participants. They also displayed certain biases, such as expressing more trust in others and showing less concern about technological threats. Compared with their human surrogates, the twins also appeared more rational.
AI researcher and economist Hadi Hosseini of Penn State University has seen a similar effect in his own research into how AI agents make health care decisions when resources are scarce: “LLMs distort human judgment in a very specific direction toward something that is very rational [and] more reasonable than actually what people are.”
But he says there are ways that the digital twins could be improved. The researchers used “a very static set of questions,” he says. Adding more dynamic approaches, such as having a chatbot shadow an individual across the day or converse regularly with a person, as is already common, could make for a better dataset, he says.
Toubia wants to try more complex training methods. Still, he says, there are cases where even digital twins of the type tested in this study would be useful. For instance, sometimes researchers need a detailed response to a question. Tired humans might provide single-sentence answers while indefatigable digital twins churn out essays. Similarly, pretesting an experiment with digital twins could help researchers ensure their design is working before sapping their human respondents’ limited bandwidth.
But Toubia also urges humility. Social scientists tend to think that scales and surveys capture the full range of human experience, he says. But it’s incredibly hard to predict human behavior with a machine. “We need to be realistic in terms of the expectations we have from synthetic data.”
文章标题:人工智能代理尚未准备好取代人类在行为研究中的角色
文章链接:https://news.qimuai.cn/?post=4975
本站文章均为原创,未经授权请勿用于任何商业用途