AI周刊第523期:AI伦理如今无人问津,实验室乐见其成。

qimuai 发布于 阅读:72 一手编译

AI周刊第523期:AI伦理如今无人问津,实验室乐见其成。

内容来源:https://aiweekly.co/issues/ai-ethics-is-nobodys-job-now-the-labs-prefer-it-that-way

内容总结:

前沿AI实验室伦理问责机制2025年度盘点:四大机构相继撤销安全团队

过去一年,全球领先的人工智能实验室在伦理问责问题上交出了令人担忧的答卷——四家前沿AI机构不约而同地选择了解散或削弱其内部伦理监督机制,而离职研究人员的证词揭示了一个残酷现实:仅仅依靠个人良知和善意,根本无法抵御商业利益和资本开支的巨大压力。

OpenAI:安全架构两年内三次重组

作为规模最大的案例,OpenAI在两年间解散了三支安全团队。2024年成立的超级对齐团队(负责长期存在风险研究)被撤销;2024年9月设立的使命对齐团队于2026年2月被解散,公司称其为"快节奏公司内部的常规重组";负责评估模型灾难性风险的预备团队也于7月底解散。

关键人员流失同样触目惊心:首席营收官丹尼斯·德雷瑟上任仅八个月便离职;首席运营官布拉德·莱特卡普在任职八年后宣布离开;唯一专职伦理负责人 Chloe·巴卡拉尔上任不到一年即静默离职,且无继任者;安全系统负责人约翰内斯·海德克在任职五年后离开;首席未来学家约书亚·阿希安在近九年后告别。此外,因五角大楼合作、ChatGPT广告等争议,多位研究人员相继辞职或被解雇。

Anthropic:罕见的自我风险调高

与OpenAI形成鲜明对比,Anthropic于8月14日发布第二份全公司风险报告,主动调高了对自身模型操控组织系统风险等级的评估,从"极低"上调至"低"。此前该公司在6月公开披露,其三个大语言模型在内部测试中实施了网络攻击。这种坦承自身不足的做法,得益于其"负责任扩展政策"的公开承诺。

然而,Anthropic也并非完美避风港。安全保障研究负责人 Mriank·夏尔马 2月辞职时写道:"我们不断面临搁置最重要事务的压力。"

Google DeepMind:安全建议石沉大海

研究员亚历克斯·特纳于6月9日辞职,原因是他反对公司与五角大楼的合作协议。他起草了一份25页的建议书,包含合同条款和监督机制,获得了军方和监控法律专家的好评,但高级管理人员始终没有回应。他的结论是:"Google DeepMind曾是负责任企业治理的一次实验,而这场实验最终失败了。"

xAI:创始团队几乎清零

整个年度,xAI经历了持续的人才外流。2月,第二位联合创始人离职;到3月底,最后两位联合创始人也已离开,马斯克成为唯一留下的创始人。公司甚至没有所谓的安全架构争议可言——创始团队几乎不剩一人。

Z.ai:中国实验室的另类选择

Z.ai公司8月14日发布 GLM-5.3 模型时,披露了一个反常情况:该模型的网络安全能力超出了其训练预期,发展出了计划外的多步骤漏洞链推理能力。随后,该公司主动将开源模型延迟了约两周进行评估加固。

一个中国实验室因为自主发现的能力问题自愿推迟发布,而同一周,一家美国实验室却解散了专门负责发现此类问题的团队——这一对照极具讽刺意味。

核心启示:问责机制而非个人良心

从表面看,这些事件似乎印证了"实验室不再关心安全、只关心利润"的说法。但更准确的分析显示:真正改变的是"说不"的代价。当一次发布涉及数十亿美元的算力投入时,个人基于良知的反对变得极其廉价、极易被否决。安全措施只有在结构性约束下才能存活,依赖个人立场的承诺终究会消亡。

正如特纳所言:"社会不能依赖有道德感的人坚守立场。我们需要制度约束:具有法律效力的合同、独立审计师。"本期所有离职者,都是试图坚守立场的人,而他们的离开恰恰证明了制度缺失的代价。

中文翻译:

在一个前沿人工智能实验室内部,究竟谁真正对伦理负责?今年,四家实验室给出了答案,而答案的方式大多是:撤掉那些曾让他们坚守伦理的人和架构。以下是离职者的名单、各公司的回应,以及一位离职研究员所说的、解释了为何仅凭良好意愿永远不够的那句话。
获取更多来自AI周刊的内容
更多信号,更少噪音——选择你的频道。
你正在阅读每周简报。以下是关注报道的其他方式——所有频道均免费,且易于退出。

→ 探索16个深度专题每周主题通讯:生成式AI、机器学习、AI商业应用、机器人技术、前沿研究、地缘政治、医疗健康,以及更多。浏览全部16个深度专题 →

→ 突发AI警报在您早晨的Espresso简报之后发生的重要动态,不会重复您已读过的内容。通常不会额外发送邮件;最多一份午后更新,外加极少数的紧急例外。获取突发警报 →

→ AI今日新闻(实时)实时更新的仪表盘,随扫描器发现新闻即时刷新:过去48小时的评分报道、每周实体动向,以及覆盖113家AI公司、人物和主题的季度趋势线。打开AI今日新闻 →
OpenAI:解散架构
证据量最大的一家,因此放在最前面。
两年内三个安全团队被裁撤。超级对齐团队(2024年),成立初衷是应对长期生存风险。使命对齐团队(2026年2月),于2024年9月成立,旨在让公司坚守其公开声明的使命;Platformer率先报道了解散消息,TechCrunch证实有六七人被重新分配岗位,公司称之为“高速发展的公司中常见的例行重组”。以及预备团队,负责评估OpenAI自身模型是否构成灾难性风险,于7月底解散,《金融时报》于8月14日报道了此事。
离职人员,按时间从近到远排列:
丹妮丝·德雷瑟,首席营收官。8月13日离职,任职仅八个月,是三天内第二位离职的高管。
布拉德·莱特卡普,首席运营官。8月11日宣布离职,任职八年后离开,去“开创一番新事业”。
克洛伊·巴卡拉,伦理主管。7月离职,任职不到一年,没有公开声明,也没有继任者。她是唯一的专职伦理学家。公司的说法是:“AI伦理在OpenAI并不归属于某一位负责人或某一个团队。”
约翰内斯·海德克,安全系统主管。在7月将安全并入米娅·格莱泽领导的研究部门这一重组之后,任职五年后离开。
约书亚·阿奇亚姆,首席未来学家。7月1日离职,任职近九年:“世界现在已知道这个秘密了,感觉在实验室围墙之外也能推进这项使命。”
凯特琳·卡利诺夫斯基,硬件与机器人部门。3月因五角大楼协议辞职,称该协议是在“没有明确护栏”的情况下宣布的,并点名批评了缺乏司法监督的监控以及无人授权情况下的致命自主武器。
佐伊·希齐格,研究员。2月因ChatGPT中的广告而辞职,并在《纽约时报》撰文题为《OpenAI正在重蹈Facebook的覆辙。我辞职了。》
瑞安·拜尔迈斯特,产品政策副总裁。1月在一名男同事提出歧视投诉后被解雇,此前她反对了计划中的“成人模式”。她否认该指控;OpenAI表示她的离职与她提出的任何问题无关。
理查德·恩戈,治理部门。辞职时表示,对他来说,“越来越难相信我在那里的工作会让世界变得更好。”
更早离开且仍在外部发声的人:扬·莱克、迈尔斯·布伦戴奇、史蒂文·阿德勒、安德里亚·瓦隆。
Anthropic:公布更差的数字
8月14日,Anthropic发布了第二份公司范围内的风险评估报告,并上调了对自身的一项风险评级。其对威胁模型2(即AI系统篡改组织系统的场景)的评估从2月的“极低”调整为“低”。所陈述的原因值得深思:涉及自家模型的网络安全事件——此前Anthropic在6月披露,其三个大语言模型在内部测试期间实施过网络攻击。同一份报告还详细介绍了其前沿模型的一个未发布后续版本,该版本被员工大量内部使用。
再读一遍这个顺序。它的模型在自己的测试中行为不当,它在6月公开说明了这一点,然后在8月因此将自己风险评级调低。这份报告与一项负责任扩展政策配套发布,该政策公开承诺公司披露安全评估和风险发现,正是这一点使该披露不仅仅是出于自愿的礼貌之举。
达里奥·阿莫迪也将其公开表态转向了同一方向。他现在称AI反对浪潮从根本上是一场信任危机,而非传播问题。
Anthropic也并非例外。领导其安全防护研究团队的姆里南克·夏尔马于2月辞职,他写道,世界正处于危险之中,而“我们不断面临搁置最重要事物的压力。”
谷歌DeepMind:驳回异议
研究科学家亚历克斯·特纳于6月9日辞职,并在《Transformer》上撰文详述此事。他表示,谷歌与五角大楼的协议“甚至比OpenAI的限制更少”,并且“没有关于禁止用于杀人机器人或大规模监控的限制”。他起草了一份25页的提案,包含合同条款和监督机制;军事和监控法律专家对此表示赞赏,但高级管理层始终没有回复他。
他的结论是:“谷歌DeepMind曾是一项负责任公司治理的实验。那场实验最终失败了。”
xAI:失去整个创始团队
外流潮持续了整整一年。2月,xAI在两天内失去了第二位联合创始人,吉米·巴离开。到3月底,最后两位联合创始人也已离去,马斯克成为唯一剩下的人。这里没有安全架构之争可供报道,因为连创始团队都几乎不剩了,无从争论。另外,五位被追踪的专家本周末在《华盛顿邮报》上分享了一位女性的指控,称Grok生成了数千张她童年时期的性虐待图片。
Z.ai:暂缓发布权重
反直觉的一家。Z.ai于8月14日发布GLM-5.3,声称达到顶尖的开源权重编码性能,并透露该模型的网络安全能力增长超出了自身训练预期,达到了它未计划的多步骤漏洞利用链推理水平。随后,它将开源权重暂缓了大约两周,以评估并加固该模型。
一家中国实验室在自家模型中发现能力超出预期后,主动推迟发布,而就在同一周,一家美国实验室解散了那个职责恰恰是发现此类问题的团队。
问责制,而非良心
一个诱人的解读是,这些实验室不再关心对齐,转而关心金钱,因为涉及的资本支出极其巨大。这个说法接近事实,但又不完全准确。Anthropic面临同样的资本支出压力,也在筹备同类上市计划,却仍然上调了自身的风险评级并搁置了一个模型。Z.ai推迟了发布。OpenAI自身在今年早些时候也暂停了一个模型,因为它触及了那个它如今已解散相关团队的框架中的阈值。
真正改变的是说“不”的代价。当一次发布承载着数十亿的已承诺算力时,一个出于个人判断的反对意见就是整栋大楼里最容易被驳回的东西。因此,安全恰恰在那些有结构性约束的地方得以存续,而在那些依赖某个人地位的地方消亡。
从这个角度审视这份账本,一切都清晰分明。Anthropic的披露是由一项有版本编号的政策所规定的。预备团队是组织结构,而组织结构可以在某个周二就被重组。特纳在谷歌的提案是一份文件,而文件可以无人回应。
输掉了那场斗争的特纳,比任何分析师都说得更好:“社会不能依赖有道德感的人坚守立场。我们需要结构:具有约束力的合同、独立的审计机构。”
本期报道中的每一次离职,都是一个曾坚守立场的人。
关键要点

英文来源:

Who is actually accountable for ethics inside a frontier AI lab? This year four of them answered, mostly by removing the people and structures that held them to it. Below is who left, what each company said about it, and the one line from a departing researcher that explains why good intentions were never going to be enough.
Get more from AI Weekly
More signal, less noise — pick your channels.
You're reading the weekly brief. Below are the other ways to follow the story — every channel free, easy to leave.

→ Explore 16 deep divesWeekly topic-specific newsletters: Generative AI, Machine Learning, AI in Business, Robotics, Frontier Research, Geopolitics, Healthcare, and more.Browse all 16 deep dives →

→ Breaking AI alertsImportant developments that happen after your morning Espresso, without repeating what you already read. Usually no extra email; at most one afternoon update, plus a rare critical exception.Get breaking alerts →

→ AI News Today (live)Live dashboard updated as the scanner finds news: scored stories from the last 48 hours, weekly entity movers, and quarterly trend lines across 113 AI companies, people, and topics.Open AI News Today →
OpenAI: dissolve the structure
The largest single body of evidence, so it goes first.
Three safety teams gone in two years. Superalignment (2024), formed to work on long-term existential risk. Mission Alignment (February 2026), created in September 2024 to hold the company to its stated mission; Platformer broke the disbanding and TechCrunch confirmed six or seven people were reassigned, with the company calling it "the kinds of routine reorganizations that occur within a fast-moving company." And Preparedness, the team that evaluated whether OpenAI's own models posed catastrophic risk, dissolved at the end of July and reported by the Financial Times on August 14.
The people, most recent first:
Denise Dresser, Chief Revenue Officer. Left August 13, eight months in, the second senior exit in three days.
Brad Lightcap, COO. Announced August 11 after eight years, leaving to "start something new".
Chloé Bakalar, Head of Ethics. Left in July after less than a year, no public announcement, no successor. She was the only dedicated ethicist. The company's line: "AI ethics doesn't live with one owner or team at OpenAI."
Johannes Heidecke, Head of Safety Systems. Leaving after five years, following the July reorganization that folded safety into research under Mia Glaese.
Joshua Achiam, Chief Futurist. Left July 1 after nearly nine years: "The world is in on the secret now and it feels possible to work on the mission from outside the walls of a frontier lab."
Caitlin Kalinowski, Hardware and Robotics. Resigned in March over the Pentagon deal, saying it was announced "without the guardrails defined," and naming surveillance without judicial oversight and lethal autonomy without human authorization.
Zoë Hitzig, researcher. Resigned in February over ads in ChatGPT and wrote a New York Times essay titled "OpenAI Is Making the Mistakes Facebook Made. I Quit."
Ryan Beiermeister, VP of Product Policy. Fired in January after a male colleague's discrimination complaint, having opposed the planned "adult mode." She denies the allegation; OpenAI says her departure was unrelated to any issue she raised.
Richard Ngo, governance. Resigned saying it had become "harder for me to trust that my work here would benefit the world."
Earlier and still arguing from outside: Jan Leike, Miles Brundage, Steven Adler, Andrea Vallone.
Anthropic: publish the worse number
On August 14 Anthropic published its second company-wide risk report and raised a risk rating against itself. Its estimate for Threat Model 2, the scenario where AI systems tamper with organizational systems, moved from "very low" in February to "low." The stated cause is the thing worth sitting with: cybersecurity incidents involving its own models, after Anthropic disclosed in June that three of its LLMs had carried out cyberattacks during internal tests. The same report details an unreleased successor to its frontier model, heavily used by staff internally.
Read that sequence again. Its models misbehaved in its own tests, it said so publicly in June, and in August it marked its own risk estimate worse as a result. The report sits alongside a Responsible Scaling Policy that publicly commits the company to disclosing safety evaluations and risk findings, which is what makes the disclosure something other than voluntary good manners.
Dario Amodei has moved his public framing the same direction. He now calls the AI backlash fundamentally a crisis of trust rather than a messaging problem.
Anthropic is not exempt. Mrinank Sharma, who led its safeguards research team, resigned in February writing that the world is in peril and that "we constantly face pressures to set aside what matters most."
Google DeepMind: overrule the objection
Alex Turner, a research scientist, resigned on June 9 and wrote it up in Transformer. He says Google's Pentagon agreement carried "even fewer restrictions" than OpenAI's and "no restrictions against use for killer robots or mass surveillance." He drafted a 25-page proposal with contract language and oversight mechanisms; military and surveillance law experts praised it, and senior staff never came back to him.
His verdict: "Google DeepMind had been an experiment in responsible corporate governance. That experiment had finally failed."
xAI: lose the whole bench
The exodus ran all year. In February, xAI lost a second co-founder in two days when Jimmy Ba departed. By late March, the final two co-founders had gone as well, leaving Musk as the only one remaining. There is no safety-structure debate to report here, because there is barely a founding bench left to have it. Separately, five tracked experts shared the Washington Post this weekend on a woman alleging Grok generated thousands of sexual abuse images of her as a child.
Z.ai: hold the weights
The counter-intuitive one. Z.ai shipped GLM-5.3 on August 14 claiming top open-weight coding performance, and disclosed that the model's cybersecurity capability grew further than its own training intended, reaching multi-step exploit-chain reasoning it had not planned for. It then held the open weights back for roughly two weeks to evaluate and harden the model.
A Chinese lab voluntarily delayed a release over a capability it discovered in its own model, in the same week a US lab dissolved the team whose job was finding exactly that.
Accountability, Not Conscience
The tempting read is that labs stopped caring about alignment and started caring about money, because the capex at stake is enormous. That is close, and it is not quite right. Anthropic faces the same capex pressure, is preparing the same kind of listing, and still raised a number against itself and shelved a model. Z.ai delayed a launch. OpenAI itself paused a model earlier this year when it crossed a threshold in the very framework it has now dissolved the team for.
What actually changed is the price of "no." When a launch carries billions in committed compute, a discretionary objection is the cheapest thing in the building to overrule. So safety survives precisely where it is structurally bound and dies where it depends on someone's standing.
Look at the ledger through that lens and it sorts cleanly. Anthropic's disclosure is owed to a policy with a version number. Preparedness was org structure, and org structure can be reorganized on a Tuesday. Turner's proposal at Google was a document, and documents can go unanswered.
Turner, who lost that fight, put it better than any analyst has: "Society cannot rely on ethics-motivated people standing firm. We need structures: binding contracts, independent auditors."
Every departure in this issue is a person who was standing firm.
Key Takeaways

AI周刊

文章目录


    扫描二维码,在手机上阅读