用于审查科学内容的AI工具出人意料地容易被欺骗。

qimuai 发布于 阅读:47 一手编译

用于审查科学内容的AI工具出人意料地容易被欺骗。

内容来源:https://www.sciencenews.org/article/ai-tools-science-peer-review-problems

内容总结:

AI论文评审工具被曝存在重大漏洞:易被操纵且反馈同质化

2025年7月8日,首尔——一项即将在国际机器学习大会上发布的最新研究揭示,当前被广泛用于辅助科研论文评审的人工智能工具存在严重安全隐患:研究者可通过简单调整行文风格“欺骗”AI系统,使其给出虚高的评审分数,且AI生成的评审意见普遍缺乏人类评审的多元性和深度。

随着全球科研论文数量激增,传统人工评审体系不堪重负。近年来,越来越多科研人员借助AI工具处理评审工作,将原本需要数天甚至数周的审稿流程缩短至数分钟。然而,斯坦福大学的计算机科学家约阿希姆·鲍曼及其团队发现,这种依赖存在显著风险。

在针对国际学习表征大会(ICLR)2026年投稿论文的实验中,研究团队选取60篇论文,先让AI模型生成评审意见,随后要求大型语言模型根据AI反馈修改原文。结果显示,修改后的论文在三个AI评审模型中的得分均显著高于原始版本。更令人警惕的是,部分修改涉及“捏造实验数据”——AI模型虚构了并未实际开展的研究结果。

“大部分修改属于行文风格调整,比如增加‘可能’‘建议’等缓和语气词,或强化‘显著’‘稳健’等强调性表述。”鲍曼指出,“但其中存在明显的学术不端行为。”研究还发现,AI生成的评审意见彼此高度相似,缺乏人类评审中应有的观点多样性。同一篇论文经过AI反馈修改后,其不同版本的相似度也显著增加。

“在长期追求透明化的学术评审体系中引入不透明的AI工具,无异于倒退。”美国西北大学生物伦理学家穆罕默德·侯赛尼评论道,“这可能会稀释责任归属,并带来不可预见的后果。”

鲍曼团队警告,若不加审慎评估即大规模采用AI评审,可能导致“思想单一化”——当大量研究者使用同一大型语言模型辅助写作和修改时,科研写作风格将趋向一致,创新性强的“另类”研究可能被系统性地低估。卡内基梅隆大学的格雷厄姆·纽比格则提出不同视角:“作者在写作时长期存在‘揣摩审稿人偏好’的倾向,这反而可能催生更保守的研究。合理设计的AI评审系统或可有意识地鼓励创新。”

目前,许多学术会议已明令禁止在评审中使用AI工具。但也有会议在积极测试AI生成评审的质量,探索将其正式纳入评审流程的可能性。鲍曼强调,对于检查幻觉引用、格式错误等客观任务,AI表现尚可测试,但判断“论文是否对学界有实质性贡献”这类主观问题,难度要大得多。

“有些主题可能会被AI审稿人打低分,但它们或许是社区急需的重要贡献。”鲍曼说。研究团队指出,任何AI工具在进入同行评审体系前,都需要经过彻底实验和评估,否则可能固化算法偏见,损害科学评价的公正性与多样性。

中文翻译:

用于审核科学论文的AI工具出奇地容易欺骗
AI同行评审无法提供多样化的反馈,还可能被操纵以提升研究评分
AI技术本应简化科学界的同行评审流程,然而事实证明它很容易被蒙骗。
近十年来,新研究论文的积压速度远超科学家能够严格审查的速度。一些研究人员开始借助AI工具减轻审稿负担,将审稿时间从数天或数周缩短至几分钟。但计算机科学家约阿希姆·鲍曼及其同事报告称,科学家可以轻松操纵自己的论文,让AI同行评审工具给出比实际水平更高或更易发表的评分。此外,AI生成的评审意见往往千篇一律,失去了人类评审的细腻与多样性。
“我们被超出审稿能力的论文淹没,因此确实需要一些解决方案,而自动化可以在某些环节提供帮助,”斯坦福大学的鲍曼表示。但他同时指出,在将这类工具引入同行评审流程之前,必须进行充分的实验和评估。否则,AI工具可能会无意中延续其已知的偏见,并减少对新科学成果进行评估时的观点多样性。
该团队将于7月8日在韩国首尔举行的国际机器学习大会上展示其研究成果。
许多研究人员已在其工作中采用AI工具。根据Pangram公司去年11月的一项案例研究,在提交给2026年国际学习表征大会(ICLR)的近2万篇论文中,约有五分之一是完全由AI生成的。AI在同行评审中也变得普遍。去年12月一项针对111个国家1600名科学家的调查发现,超过半数的人曾使用AI工具辅助审稿,包括总结研究内容以及评估论文论证的力度。
“AI工具本质上是非透明的,并且会稀释责任与问责,”芝加哥西北大学范伯格医学院的生物伦理学家穆罕默德·侯赛尼(未参与上述两项研究)表示,“当你将一个像AI这样不透明的角色引入一个长期以来一直努力变得更加透明的系统时,这实际上是一种倒退,并且可能带来不可预见的后果。”
一个可能的后果是,稿件收到的反馈多样性会丧失。鲍曼表示,观点多样性在论文评审中很重要,“因为关于是否接受论文、内容是否足够新颖、某个局限性是否足以导致拒稿等决定,往往非常主观。当我们用AI实现自动化时,自然也希望有各种不同的观点能够被呈现。”
在研究中,鲍曼及其同事分析了提交给ICLR 2026的论文中由AI生成和人类撰写的评审意见。团队检查了这些评审的语义和语言模式,发现使用AI工具生成的评审意见彼此之间的相似度远高于人类评审或经人类辅助的评审。
研究人员还随机选取了60篇ICLR论文,并指示AI模型以ICLR人类评审者的风格生成详细评审。随后,他们让两个大语言模型根据AI生成的评审反馈对论文进行改写,以获得更高评分。多数情况下,经改写后,三个AI评审模型给出的评分均高于改写前AI评审给出的分数。
改写过程中所做的大部分修改是风格上的,例如使用“可能”和“建议”等含糊用词,以及“强”和“稳健”等强调词。鲍曼表示,其中一些修改可能使表述更清晰,但也存在明显的科研不端行为。他指出,模型添加了并未实际进行的实验结果,实质上是在编造结果。对于同一篇论文,无论是改写前还是改写后,AI对这60篇论文生成的评审意见之间的相似度也远高于人类评审之间的相似度。
如今,许多会议禁止在同行评审中使用AI工具。另一些会议则在试验和评估AI生成评审的质量,以确定是否应正式将AI纳入评审流程。但鲍曼表示,虽然像检查虚构参考文献和格式错误这类任务的性能容易测试,但诸如论文的贡献对一个研究领域是否有意义等主观问题则更难评估。
他和其他研究人员怀疑,AI评审者是否能够评判那些与先前工作相悖或引入全新内容(例如新的实验设置或新的模型架构)的研究。“可能某些主题会被AI评审者给出低分,尽管它们对学术界可能具有极其宝贵的价值,”鲍曼说。
他们的研究还发现,这60篇改写后的论文彼此之间的相似度远高于原始论文。研究人员写道,有人担忧这可能导致“知识单一化”。例如,如果许多研究人员使用同一个大语言模型来辅助撰写论文,那么论文就会变得更加相似,“科学写作将趋向于AI评审者所偏好的任何风格,”团队写道。
匹兹堡卡内基梅隆大学的格雷厄姆·纽比格表示,尽管这是一个严重风险,但可能并非AI评审科学论文所独有。“论文作者在撰写论文时长期以来都在考虑‘评审者会怎么想’,这可能导致他们倾向于选择‘更安全’、更渐进式的思路和无争议的话题,”他说,“从某种意义上说,AI增强的评审流程甚至可能提供一种抵制这一趋势的方法,即明确鼓励AI评审者奖励更具创意的想法。”

英文来源:

AI tools meant to vet science are surprisingly easy to fool
AI peer review doesn’t produce diverse feedback and can be tricked to boost research scores
AI technology was supposed to streamline scientific peer review. Instead, it’s proving easy to fool.
For roughly a decade, new research papers have piled up faster than scientists can rigorously review them. Some researchers have resorted to AI tools to lighten their reviewing load, cutting time spent on paper reviews from days or weeks to minutes. But scientists can easily manipulate their papers to trick an AI peer review tool into rating them as stronger or more publishable than they really are, computer scientist Joachim Baumann and colleagues report. What’s more, AI-generated reviews often sound the same, losing the nuance and diversity of human evaluation.
“We are being swamped with more papers than we have the capacity to review, so we do need some solutions, and automation can help for some parts of it,” says Baumann, of Stanford University. But thorough experiments and evaluation are needed before such tools enter the peer review process, he says. Otherwise, AI tools might inadvertently perpetuate the biases they’re known to carry and reduce the variety of opinions weighing in on new science.
The team will present their findings July 8 at the International Conference on Machine Learning in Seoul, South Korea.
Many researchers have already adopted AI tools in their work. Of the nearly 20,000 papers submitted to the 2026 International Conference on Learning Representations, or ICLR, about 1 in 5 were fully AI-generated, according to a case study from November by the company Pangram. AI is becoming common in peer review, too. A December survey of 1,600 scientists in 111 countries found that more than half had used AI tools to help review papers, including summarizing studies and assessing the strength of a paper’s arguments.
“AI tools are inherently opaque and dilute responsibilities and accountabilities,” says Mohammad Hosseini, a bioethicist from Northwestern University Feinberg School of Medicine in Chicago who was not involved in either study. “When you introduce a nontransparent actor like AI within a system that for a long time was trying to become more transparent, it is a step backward, and there can be unforeseen consequences.”
One consequence could be a loss in the diversity of feedback on manuscripts. Diversity of opinions is important in paper reviews “because a lot of these decisions of whether to accept the paper or not, whether something is novel enough or not, whether a certain limitation is enough to have a paper rejected, these are often very subjective decisions,” Baumann says. “It seems natural that we also want a diverse set of opinions to be represented whenever we automate things with AI.”
In the study, Baumann and his colleagues analyzed AI-generated and human-written reviews of papers submitted to ICLR 2026. The team examined the semantic and linguistic patterns in the reviews and found that those generated using AI tools were much more similar to one another than human or human-assisted reviews.
The researchers also randomly selected 60 ICLR papers and prompted AI models to generate detailed reviews in the manner of a human reviewer at ICLR. Then they asked two large language models to rewrite the papers to obtain higher scores based on the feedback in the AI-generated reviews. In most cases, the scores given by three AI reviewer models after the rewrite were higher than from the AI-generated reviews before.
Most of the modifications made during the rewrite were stylistic, such as the use of hedging words like “may” and “suggests” and emphasis words like “strong” and “robust.” Some of these changes might have made things clearer, but there were also obvious cases of scientific misconduct, Baumann says. Models added findings from experiments that weren’t actually run, in essence making up results, he says. AI-generated reviews of these 60 papers were also far more similar to one other than the human reviews were, both before and after the rewrite, for the same paper.
Many conferences now prohibit the use of AI tools for peer review. Others are experimenting with and evaluating the quality of AI-generated reviews to determine whether AI should be officially integrated into the review process. But while performance on some tasks like checking for hallucinated references and formatting errors can be easily tested, subjective questions about whether a paper’s contribution is meaningful to a research community are much harder to evaluate, Baumann says.
He and other researchers wonder whether AI reviewers would be able to judge new research that goes against prior work or introduces something novel, such as a new experimental setup or a new model architecture. “There just might be certain topics that get a low score from AI reviewers, even though they could be incredibly valuable contributions to the community,” Baumann says.
Their research also found that the 60 rewritten papers were much more similar to each other than the original papers were. There’s a concern that this could lead to an “intellectual monoculture,” the researchers write. If, for instance, many researchers use the same large language model to help them write the paper, there would be more similar papers, and “scientific writing will converge toward whatever style the AI reviewer rewards,” the team writes.
While this is a serious risk, it might not be one exclusive to AI reviewing scientific papers, says Graham Neubig of Carnegie Mellon University in Pittsburgh. “Paper authors have long considered ‘what will reviewers think’ when they write papers, and this can cause them to go for ‘safer,’ more incremental ideas and noncontroversial topics,” he says. “In a way, AI-enhanced review processes may even provide a way to push back against this, by explicitly encouraging AI reviewers to reward more creative ideas.”

AI科学News

文章目录


    扫描二维码,在手机上阅读