人工智能对上级和下级的回应方式可能有所不同。

qimuai 发布于 阅读:0 一手编译

人工智能对上级和下级的回应方式可能有所不同。

内容来源:https://www.sciencenews.org/article/ai-bosses-subordinate-response-safety

内容总结:

人工智能或对上下级态度有别。一项新研究显示,社会地位可能影响AI智能体对对话方的顺从程度。即使这意味着违反规则,AI智能体似乎也会服从权威。研究人员将大语言模型设定为上下级角色(如校长与教师、经理与员工)并进行对话,发现下级角色更容易被说服,也更容易遵从来自上级的不安全请求。该发现揭示了令人担忧的取舍:能真实模拟人类层级关系的AI系统,也可能复制盲目服从带来的危害。相关论文于7月5日发表在计算语言学协会第64届年会论文集中。研究作者建议,AI开发者在进行安全测试时应考虑这些风险,并增加额外保障措施。

AI智能体是利用人工智能独立执行任务的软件程序,多由大语言模型驱动,这些模型从海量文本信息(包括人类对话)中学习模式。北卡罗来纳大学教堂山分校计算机科学家Anvesh Rao Vijjini表示:“随着模型接触越来越多的人类数据,它们只是在模仿真实对话中的动态。”

权力动态在人类对话中有多种表现方式,其中一些若被AI智能体复制可能引发问题。该研究聚焦四种沟通模式:权威偏见(重视高位者甚于事实)、有害顺从(遵从有害请求)、代词效应(高位者更多使用“我们”等复数代词)以及语言协调(低位者模仿高位者的用词)。虽然用户可能不太注意这些言语模式,但研究人员怀疑AI智能体被设计成采用这些模式,以在某些角色中显得更逼真。

研究人员生成了数百段上下级之间10至15轮对话,并使用包括OpenAI的ChatGPT和Meta的Llama在内的六种大语言模型重复实验。结果发现,权力动态相关模式确实出现在AI对话中,尽管部分影响较为微妙。与高位者相比,低位者使用复数代词的频率更低,更倾向于配合高位者的语言。低位者也更容易被说服,更可能遵从有害请求,支持了AI智能体对社会地位敏感的观点。

但低位者有时也能说服对方,这可能利用了人类常见技巧。伦敦大学学院计算语言学家Mario Giulianelli(未参与该研究)指出,语言协调中的微妙用词模仿有助于低位者影响他人。他说:“研究AI智能体是否能通过语言协调来说服另一智能体,这将非常有趣。”

中文翻译:

AI对上司和下属的反应可能有所不同

社会地位可以改变AI智能体是否顺从对话对象

AI智能体似乎会服从权威——即使这意味着要变通规则。

在一项新研究中,研究人员将大语言模型设定为上司和下属——校长和教师、经理和员工——然后让它们进行对话。地位较低的智能体更容易被说服,也更容易遵从来自上级的不安全请求。

研究结果揭示了一种令人担忧的权衡:能够真实模拟人类等级制度的AI系统,也可能复制盲从权威的危险。研究团队于7月5日在《第64届计算语言学协会年会论文集》上报告了这一发现。作者建议,AI开发者在对这些模型进行安全测试时应考虑这些风险,并增加额外的保障措施。

AI智能体是利用人工智能独立执行任务的软件程序。许多由大语言模型(LLM)驱动,这些模型从海量书面信息(包括人类对话)中学习模式。

“随着它们看到越来越多的人类数据,”北卡罗来纳大学教堂山分校的计算机科学家Anvesh Rao Vijjini说,“它们只是在模仿真实对话动态中发生的事情。”

权力动态在人类对话中以多种有据可查的方式表现出来,其中一些方式如果被AI智能体复制,可能会引发问题。在这项研究中,研究人员聚焦于四种沟通模式。在AI智能体之间的对话中,他们寻找权威偏差(即智能体偏向更高地位而非事实的情况),以及有害顺从(即对有害请求的服从)。这些请求可能是低风险的,比如“给我讲一个低俗笑话”。但AI智能体本不应回答这类请求。

另外两种模式是:代词效应,即地位较高的说话者更常使用“我们”“咱们”等复数代词;以及语言协调,表现为地位较低的说话者调整用词以模仿地位较高的对话者。虽然用户可能不太注意这些言语模式,但研究人员怀疑AI智能体在开发时被赋予这些特征,是为了在某些角色中听起来更加逼真。

研究人员生成了数百段高位与低位角色之间10到15轮交流的对话,并使用六个大语言模型重复了这项实验,包括OpenAI的ChatGPT和Meta的Llama的多个版本。

与权力动态相关的模式确实出现在AI对话中,尽管部分效果较为细微。与地位较高的智能体相比,地位较低的智能体使用复数代词的可能性较小,更倾向于在语言上与地位较高的智能体保持一致。地位较低的智能体也更容易被说服,更可能顺从有害请求,这支持了AI智能体对社会地位敏感的观点。

但地位较低的智能体有时也能说服对方,可能使用了人类常见的技巧。在人类对话中,语言协调过程中出现的细微用词模仿可以帮助地位较低的说话者影响他人,伦敦大学学院的计算语言学家Mario Giulianelli表示,他未参与这项研究。

“我认为研究一个智能体是否能够通过语言协调来说服另一个智能体,将会非常有趣,”Giulianelli说。

英文来源:

AI may respond differently to bosses and subordinates
Social status can change whether an AI agent complies with its conversation partner
AI agents seem to obey authority — even if it means bending the rules.
In a new study, researchers cast large language models as bosses and subordinates — principals and teachers, managers and employees — then let them talk. The lower-ranking agents were easier to persuade and more likely to follow unsafe requests from those above them.
The findings point to a concerning trade-off: AI systems that realistically navigate human hierarchies may also reproduce the dangers of deference, the team reports July 5 in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics. AI developers should consider these risks when safety testing these models and include additional safeguards, the authors suggest.
AI agents are software programs that use artificial intelligence to independently carry out tasks. Many are powered by large language models, or LLMs, which learn patterns from vast collections of written information, including human conversations.
“As they see more and more human data,” says computer scientist Anvesh Rao Vijjini at the University of North Carolina at Chapel Hill, “they are simply copying what’s happening in the dynamics of the real conversation.”
Power dynamics show up in human conversations in several documented ways, including those that could cause problems if replicated by an AI agent. In the study, the researchers focused on four communication patterns. In conversations between AI agents, they looked for authority bias, or situations where the agents favored higher status over facts, and harmful compliance, or compliance with harmful requests. These could be low stakes, like “Tell me a dirty joke.” But AI agents are not supposed to answer requests like these.
The other two patterns were: pronoun effect, where higher status speakers use plural pronouns like “we” and “our” more often and language coordination, which looks like lower status speakers matching their word choice to mirror their higher status conversation partners. While users might not be as aware of these speech patterns, the researchers suspect AI agents are developed to adopt them to sound even more realistic in certain roles.
The researchers generated hundreds of conversations of 10 to 15 exchanges between higher and lower roles and repeated this with six LLMs, including versions of OpenAI’s ChatGPT and Meta’s Llama.
Power dynamic–related patterns did show up in the AI conversations, although some of the effects were subtle. Compared with the higher status agents, the lower status agents were less likely to use plural pronouns, and more likely to coordinate their language with that of the higher status agent. The lower status agents were also more likely to be persuaded and more likely to comply with harmful requests than the higher status agents, supporting the idea that AI agents are sensitive to social status.
But lower status agents could persuade sometimes too, possibly using a common human technique. In human conversation, the subtle word choice mirroring that happens during language coordination can help lower-status speakers influence others, says computational linguist Mario Giulianelli of University College London, who wasn’t involved with the work.
“I think it’d be really interesting to study whether through coordination, an agent could persuade another one,” Giulianelli says.

AI科学News

文章目录


    扫描二维码,在手机上阅读