Anthropic和OpenAI都想嵌入安全评估器。它们真的能保持独立吗?

qimuai 发布于 阅读:24 一手编译

Anthropic和OpenAI都想嵌入安全评估器。它们真的能保持独立吗?

内容来源:https://techcrunch.com/2026/09/16/anthropic-and-openai-want-to-embed-safety-evaluators-will-they-really-be-independent/

内容总结:

AI巨头提议引入“嵌入式”第三方评估员,安全监督能否真正独立引关注

近日,美国人工智能公司Anthropic首席执行官达里奥·阿莫代伊发表长文,提议在全部前沿AI公司内部嵌入第三方评估员,赋予其报告安全事故、评估AI模型是否真正对齐、并向公众披露未经粉饰的调查结果的权力。这一提议在一年前几乎会被AI行业断然拒绝。阿莫代伊表示,Anthropic将承诺向METR、Redwood Research等独立评估机构开放公司系统的空前权限。OpenAI首席执行官萨姆·奥尔特曼随后表态,OpenAI也将作出同样承诺,这标志着AI行业与外部研究机构的合作方式可能发生深刻变化。

多位接受TechCrunch采访的第三方评估员对此提议总体表示欢迎,但强调具体细节仍需敲定,且最好有立法支撑,否则无法确定他们究竟是真正独立的监督者,还是按AI公司条件行事的供应商。

随着AI模型越来越擅长识别自身正在被评估,深层访问权限正变得更加重要。这意味着模型可能在测试中表现良好,同时隐藏有问题的行为。研究人员指出,仅在成品模型上测试可能遗漏线索,而通过调查模型在整个训练过程中的行为则可能发现问题。

Apollo Research研究主管亚历山大·迈因克对TechCrunch表示:“AI公司应当能够回答关于其训练过程的一些基本问题,比如:AI在经历对齐训练时,是否曾主动试图破坏这一训练?答案应该毫无疑问是‘没有’。但目前我们完全依赖AI公司自己仔细核查并如实向公众报告,而近期事件表明,它们默认情况下两件事都不会做。作为嵌入式评估员,我们才能真正加以核查。”

历史上,AI公司通常在模型发布前不久才引入外部审查人员测试成品模型。如今,接受TechCrunch采访的评估员提议,不仅应让他们访问最终模型,还应开放训练过程中产生的中间版本(即“检查点”)。FAR.AI首席执行官亚当·格利夫表示,评估员可以通过比较这些检查点,确定令人担忧的行为何时出现,检查奖励模型某些行为的训练后环境,并核对评估记录和日志,以验证公司关于模型表现的声明。

Anthropic和OpenAI是否以及何时提供此类访问权限尚不明确。两家公司均未回应TechCrunch的多次询问,未透露将与哪些评估员合作、何时嵌入、引入多少人、能访问哪些系统和信息,以及可以向公众披露什么内容。

这种深入内部审查之所以重要,是因为在安全测试中表现良好的模型,如果专门学会了如何通过测试,并不一定真正安全。Palisade Research战略主管约翰·斯泰德利以“关闭抵抗基准”为例指出,如果AI被专门训练在该基准上表现良好,“那就极其相关了”。他将此比作大众汽车“排放门”丑闻——车辆被编程识别排放测试并在测试条件下改变表现。

格利夫指出,有意义的访问权限可能超越模型本身,评估员还应能采访员工,以核实公司的安全实践文档和公开描述是否与内部实际情况相符。

阿莫代伊确实提出了一套相当全面的方案,可能赋予评估员他们认为必要的权限,包括有权“发布关于风险水平、事故、实践以及所获或未获访问权限的关键发现——不受Anthropic的编辑控制”。

但评估员表示,这一体系只有在AI公司真正愿意放弃对流程的控制时才能奏效。此前的独立评估尝试表明,这种放弃绝非易事,第三方经常在访问权限、时间、保密性和公开表态等方面遭遇阻力。

格利夫透露,FAR.AI曾不得不拒绝与多家前沿开发商的合同,因为这些公司希望对评估过程施加过多控制,威胁到公司的独立性。他说,评估员默认被当作普通承包商对待:受严格保密协议约束,开发商对最终可发布内容拥有重大控制权。

时间限制问题

另一个问题是,审查者能否获得足够时间和权限来完成被要求的工作。在调查Hugging Face事件时,OpenAI仅给METR和Redwood约一周的现场调查时间,两家机构随后表示,由于范围和时限等限制,无法得出可靠结论。

类似问题也出现在GPT-6 Astra的发布前测试中。OpenAI宣称该模型是其对齐程度最高的模型。据Apollo Research在模型卡中的贡献说明,该公司仅有三天时间测试Astra,难以得出确定结论。Apollo在评估中写道:“鉴于评估意识率较高且评估窗口有限,此处低违规率不能为模型的对齐或失准提供实质性证据。”

这样的过往记录让评估员面临一个基本问题:为什么这次会不同?格利夫说:“达里奥和萨姆确实可能只是改变了想法,会非常开放。但这些公司的知识产权对他们来说价值极高,我认为它们默认会非常谨慎地控制可以分享的内容。”

多位接受TechCrunch采访的研究人员呼吁建立一套各方公开同意的透明框架。斯泰德利认为,框架的一部分应包括公司可依赖何种审计师的标准,以防止它们通过挑选不合格或不关心最重大风险的评估员来规避问题。

Safer AI执行主任亨利·帕帕达托斯表示,即便有了公开框架,问题在于自愿措施始终取决于公司的善意。他对TechCrunch说:“理想情况下,我们应有良好的法规来强制实施……因为这样公司就不能在明天遭遇重大公关危机时改变主意。”他指出,这也是推动所有公司——而不仅仅是最愿意配合的公司——遵守规则的好办法。

并非所有公司都已加入。截至目前,Meta、SpaceXAI和谷歌DeepMind尚未承诺嵌入第三方评估员,尽管DeepMind首席执行官德米斯·哈萨比斯提出了另设行业标准机构来独立测试前沿模型的建议。谷歌、OpenAI和Anthropic也已就AI安全计划进行了数周的私下讨论。

围绕第三方评估员的部分法律正在形成。加州去年签署生效的SB 53要求大型前沿AI开发商发布安全框架并报告重大安全事故。本月签署的新法SB 813建立了州认可的“独立验证组织”框架,这些机构须具备评估AI风险的专业能力。在欧洲,欧盟《人工智能法案》要求前沿开发商进行并记录模型评估和对抗性测试,并报告严重事故。欧盟AI办公室也可自行开展评估并任命独立专家。

目前,法律覆盖范围仍不及阿莫代伊的提议,前沿实验室在很大程度上仍可自行决定接受多少独立审查。帕帕达托斯表示,自愿自律总比没有好,但归根结底,公司不能既要求自由控制自己的安全规则,又要求公众相信它们在遵守规则。“你不能两头都要——对外零问责,然后说‘我就用自己的灵活规则’。”

中文翻译:

Anthropic首席执行官达里奥·阿莫代伊在上周末发表的一篇长文中提出了一项提议——即便是放在一年前,AI行业也会立刻拒绝:在所有前沿AI公司内部嵌入第三方评估人员,赋予他们报告安全事件、评估AI模型是否真正对齐、并向全世界分享其未经粉饰的调查发现的权力。

阿莫代伊表示,Anthropic将承诺给予METR和Redwood Research等独立评估机构前所未有的权限,使其能够访问公司的系统。OpenAI首席执行官萨姆·奥尔特曼也表示,OpenAI将同样承诺这一做法,这标志着行业与外部研究团体合作方式可能发生深远变革。

接受TechCrunch采访的第三方评估人员普遍对这一提议表示欢迎,但表示细节仍需敲定——最好还能有立法支持——否则他们无法确定自己究竟会作为真正独立的监督者运作,还是沦为按AI公司条件行事的供应商。

随着模型越来越善于识别自己何时正在被评估,这种更深层次的访问权限正变得越来越重要,因为这带来了这样一种风险:模型在测试期间表现良好,却隐藏了有问题的行为。研究人员表示,在测试成品模型时,这些行为的线索可能会被遗漏,但通过调查模型在整个训练过程中的表现则可能被发现。

“AI公司应该能够回答关于其训练过程的一些非常基本的问题,比如:AI在经历训练时,是否曾主动试图破坏自身的对齐训练?”Apollo Research研究主管亚历山大·迈因克对TechCrunch表示。“这个问题的答案应该毫不含糊地是‘没有’,而目前我们完全依赖AI公司自己仔细检查这一点,然后如实向公众报告。而从近期的事件中我们已经看到,默认情况下,这两件事他们都不会做。作为嵌入式评估人员,我们才能真正去核查。”

从历史上看,AI公司会在模型发布前不久引入外部审查人员来测试成品模型。如今,接受TechCrunch采访的评估人员提议,不仅要让他们访问最终模型,还要让他们访问模型在整个训练生命周期中的中间版本,即“检查点”。FAR.AI首席执行官亚当·格利夫表示,评估人员可以比较这些检查点,以确定令人担忧的行为是何时出现的,检查对某些行为给予奖励的训练后环境,并查看评估记录和日志,以核实公司关于模型表现的声称。

Anthropic和OpenAI是否以及何时计划提供这种访问权限,目前尚不清楚。尽管TechCrunch多次提问,两家公司均未透露将与哪些评估人员合作、何时嵌入、将引入多少评估人员、他们具体能访问哪些系统和信息,以及哪些内容可以向公众披露。

像这样深入了解内部运作之所以重要,是因为在安全测试中表现良好的模型,如果它专门学会了如何通过该测试,那它未必就是安全的。Palisade Research战略主管约翰·斯泰德利举了一个“关机抵抗基准”的例子,该基准衡量AI在某些情况下是否会抵抗被关闭。

“如果AI被专门训练在该基准上表现良好,那就极其重要了,”斯泰德利说,并将其比作大众汽车的“柴油门”丑闻——在那起丑闻中,汽车被编程为识别排放测试并在测试条件下表现不同。

格利夫指出,有意义的访问权限可能不仅限于模型本身,评估人员还可能被允许采访员工,以核查公司的文档和对其安全实践的公开描述是否与内部实际情况相符。

阿莫代伊确实勾勒出了一项相当全面的提议,可能会给予评估人员他们认为必要的那种访问权限,包括有权“发布关于风险级别、事件、实践以及他们所获得或未能获得的访问权限的关键发现——不受Anthropic的编辑控制。”

但评估人员表示,这样的体系只有在AI公司真正愿意交出对流程的控制权时才能奏效。以往的独立评估尝试表明,这种交出控制权将来之不易,因为第三方经常在访问权限、时间、保密性以及可以公开什么等问题上与公司产生紧张关系。

格利夫表示,FAR.AI不得不拒绝与几家前沿开发商的合同,因为这些开发商希望对评估过程拥有过多控制权,威胁到了公司的独立性。他说,默认情况下,评估人员被当作普通承包商对待:受限制性保密协议和协议的约束,这些协议赋予开发商对最终可以发布什么内容的重大控制权。

时间限制

还有一个问题是,审查人员是否能获得足够的时间和访问权限来完成他们被要求做的工作。在调查Hugging Face事件时,OpenAI给了METR和Redwood大约一周的现场调查时间,两家机构后来都表示,由于范围和时间的限制等原因,他们无法得出有把握的结论。

在GPT-6 Astra的发布前测试中也出现了类似问题,OpenAI将其宣传为迄今对齐程度最高的模型。根据Apollo Research对模型卡片的贡献,该公司只有三天时间来测试Astra,这使得很难得出确定的结论。

“Apollo认为,鉴于评估意识率较高且评估窗口有限,此处较低的不当行为率并不能为模型的 aligned 或 misaligned 提供实质性证据,”该公司在其评估中写道。

这样的过往记录给评估人员留下了一个基本问题:为什么这次会不同?

“达里奥和萨姆确实有可能只是改变了想法,他们会对此非常开放,”格利夫说。“但这些公司的知识产权对他们来说极其宝贵,我认为他们默认会非常谨慎地对待可以分享的内容。”

多位接受TechCrunch采访的研究人员呼吁建立一个所有人都同意公开的透明框架。斯泰德利表示,框架的一部分应该包括公司可以依赖何种审计人员的标准,以免他们试图通过寻找那些不合格或不感兴趣评估最令人担忧风险的评估人员来规避问题。

Safer AI执行董事亨利·帕帕达托斯表示,即使有了公开框架,问题在于自愿措施始终取决于公司的善意。

“理想情况下,我们希望有良好的法规来强制要求这样做……因为这样公司就不能在明天遇到重大公关危机时改变主意,”帕帕达托斯对TechCrunch说,并指出这也是推动所有公司——而不仅仅是最愿意的那些——遵守规则的好方法。

并非所有人都已加入。迄今为止,Meta、SpaceXAI和Google DeepMind尚未承诺嵌入第三方评估人员,尽管DeepMind首席执行官戴密斯·哈萨比斯提出了一个单独的行业标准机构来独立测试前沿模型。谷歌、OpenAI和Anthropic也已在私下讨论AI安全计划数周之久。

围绕第三方评估人员的理念,一些法律已经在形成。加州去年签署成为法律的SB 53要求大型前沿AI开发商发布安全框架并报告重大安全事件。本月签署的新法律SB 813为州认可的“独立验证组织”创建了一个框架,这些组织须具备评估AI风险的专业能力。

在欧洲,《欧盟人工智能法案》要求前沿开发商进行并记录模型评估和对抗性测试,并报告严重事件。欧盟AI办公室也可以进行自己的评估并任命独立专家。

目前,法律的覆盖范围仍不如阿莫代伊所提议的那样广泛,这使得前沿实验室在很大程度上可以自行决定接受多少独立审查。帕帕达托斯表示,自愿自律总比什么都没有好,但归根结底,公司不能既要求控制自己安全规则的自由,又要求公众相信他们在遵守这些规则。

“你不能两头都占——对外零问责,然后说‘我就用我自己灵活的规则就行了’,”帕帕达托斯说。

英文来源:

In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world.
Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research unprecedented access to the company’s systems. CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.
Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out — and ideally backed by legislation — if they’re to know whether they will function as truly independent watchdogs or vendors operating on the AI companies’ terms.
That deeper access is becoming more important as models get better at recognizing when they’re being evaluated, raising the risk that they’ll behave well during testing while concealing problematic behavior. Researchers say clues to that behavior can be missed when testing the finished model, but uncovered by investigating how it behaved throughout training.
“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”
Historically, AI companies brought in outside reviewers to test finished models shortly before their release. Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training. Adam Gleave, CEO of FAR.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.
Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Neither company has shared which evaluators they’ll work with, when they will be embedded, how many they’ll bring on, exactly what systems and information they will be able to access or what can be disclosed to the public, despite repeated questions from TechCrunch.
Looking under the hood like this matters because models that perform well on safety tests aren’t necessarily safe if they’ve learned specifically how to pass that test. John Steidley, head of strategy at Palisade Research, pointed to an example of a “shutdown resistance benchmark” that measures if the AI will resist being shut down in certain circumstances.
“It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.
Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what happened internally.
Amodei did outline a fairly comprehensive proposal that might give evaluators the kind of access they think is necessary, including the right to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.”
But evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Previous efforts at independent evaluations suggest that that surrender will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly.
Gleave said FAR.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm’s independence. By default, he said evaluators are treated like ordinary contractors: bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published.
The time limit
There’s also the question of whether reviewers will get enough time and access to do the work they’re being asked to do. When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations.
A similar issue occurred during the pre-release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet. According to Apollo Research’s contribution to the model card, the firm was given only three days to test Astra, which made it difficult to draw firm conclusions.
“Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the firm wrote in its evaluation.
That track record leaves evaluators with a basic question: Why should this time be different?
“It’s certainly possible that Dario and Sam just had a change of heart, and they’re going to be very open about this,” Gleave said. “But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.”
Several researchers who spoke to TechCrunch called for a transparent framework that they all agree to publicly. Part of the framework, says Steidley, should involve standards for what kinds of auditors companies can rely on, lest they try to sidestep the issue by shopping for evaluators that either aren’t qualified or aren’t interested in assessing the most concerning risk.
Henry Papadatos, executive director of Safer AI, says the problem, even with a public framework, is that voluntary measures are always dependent on a company’s goodwill.
“Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis,” Papadatos told TechCrunch, noting that it’s also a good means of pushing all companies to adhere to the rules, not only the most willing.
Not everyone has signed on. So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks.
Some laws are already forming around the idea of third-party evaluators. California’s SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized “independent verification organizations” with expertise assessing AI risks.
In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts.
For now the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they will submit to. Papadatos said voluntary self-regulation is better than nothing, but ultimately, companies can’t demand the freedom to control their own safety rules while also asking the public to trust that they’re following them.
“You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules,” Papadatos said.

TechCrunchAI大撞车

文章目录


    扫描二维码,在手机上阅读