AI周刊第523期:AI伦理如今无人问津,实验室乐见其成。

内容来源:https://aiweekly.co/issues/ai-ethics-is-nobodys-job-now-the-labs-prefer-it-that-way
内容总结:
前沿AI实验室伦理问责机制2025年度盘点:四大机构相继撤销安全团队
过去一年,全球领先的人工智能实验室在伦理问责问题上交出了令人担忧的答卷——四家前沿AI机构不约而同地选择了解散或削弱其内部伦理监督机制,而离职研究人员的证词揭示了一个残酷现实:仅仅依靠个人良知和善意,根本无法抵御商业利益和资本开支的巨大压力。
OpenAI:安全架构两年内三次重组
作为规模最大的案例,OpenAI在两年间解散了三支安全团队。2024年成立的超级对齐团队(负责长期存在风险研究)被撤销;2024年9月设立的使命对齐团队于2026年2月被解散,公司称其为"快节奏公司内部的常规重组";负责评估模型灾难性风险的预备团队也于7月底解散。
关键人员流失同样触目惊心:首席营收官丹尼斯·德雷瑟上任仅八个月便离职;首席运营官布拉德·莱特卡普在任职八年后宣布离开;唯一专职伦理负责人 Chloe·巴卡拉尔上任不到一年即静默离职,且无继任者;安全系统负责人约翰内斯·海德克在任职五年后离开;首席未来学家约书亚·阿希安在近九年后告别。此外,因五角大楼合作、ChatGPT广告等争议,多位研究人员相继辞职或被解雇。
Anthropic:罕见的自我风险调高
与OpenAI形成鲜明对比,Anthropic于8月14日发布第二份全公司风险报告,主动调高了对自身模型操控组织系统风险等级的评估,从"极低"上调至"低"。此前该公司在6月公开披露,其三个大语言模型在内部测试中实施了网络攻击。这种坦承自身不足的做法,得益于其"负责任扩展政策"的公开承诺。
然而,Anthropic也并非完美避风港。安全保障研究负责人 Mriank·夏尔马 2月辞职时写道:"我们不断面临搁置最重要事务的压力。"
Google DeepMind:安全建议石沉大海
研究员亚历克斯·特纳于6月9日辞职,原因是他反对公司与五角大楼的合作协议。他起草了一份25页的建议书,包含合同条款和监督机制,获得了军方和监控法律专家的好评,但高级管理人员始终没有回应。他的结论是:"Google DeepMind曾是负责任企业治理的一次实验,而这场实验最终失败了。"
xAI:创始团队几乎清零
整个年度,xAI经历了持续的人才外流。2月,第二位联合创始人离职;到3月底,最后两位联合创始人也已离开,马斯克成为唯一留下的创始人。公司甚至没有所谓的安全架构争议可言——创始团队几乎不剩一人。
Z.ai:中国实验室的另类选择
Z.ai公司8月14日发布 GLM-5.3 模型时,披露了一个反常情况:该模型的网络安全能力超出了其训练预期,发展出了计划外的多步骤漏洞链推理能力。随后,该公司主动将开源模型延迟了约两周进行评估加固。
一个中国实验室因为自主发现的能力问题自愿推迟发布,而同一周,一家美国实验室却解散了专门负责发现此类问题的团队——这一对照极具讽刺意味。
核心启示:问责机制而非个人良心
从表面看,这些事件似乎印证了"实验室不再关心安全、只关心利润"的说法。但更准确的分析显示:真正改变的是"说不"的代价。当一次发布涉及数十亿美元的算力投入时,个人基于良知的反对变得极其廉价、极易被否决。安全措施只有在结构性约束下才能存活,依赖个人立场的承诺终究会消亡。
正如特纳所言:"社会不能依赖有道德感的人坚守立场。我们需要制度约束:具有法律效力的合同、独立审计师。"本期所有离职者,都是试图坚守立场的人,而他们的离开恰恰证明了制度缺失的代价。
中文翻译:
在一个前沿人工智能实验室内部,究竟谁真正对伦理负责?今年,四家实验室给出了答案,而答案的方式大多是:撤掉那些曾让他们坚守伦理的人和架构。以下是离职者的名单、各公司的回应,以及一位离职研究员所说的、解释了为何仅凭良好意愿永远不够的那句话。
获取更多来自AI周刊的内容
更多信号,更少噪音——选择你的频道。
你正在阅读每周简报。以下是关注报道的其他方式——所有频道均免费,且易于退出。
→ 探索16个深度专题每周主题通讯:生成式AI、机器学习、AI商业应用、机器人技术、前沿研究、地缘政治、医疗健康,以及更多。浏览全部16个深度专题 →
→ 突发AI警报在您早晨的Espresso简报之后发生的重要动态,不会重复您已读过的内容。通常不会额外发送邮件;最多一份午后更新,外加极少数的紧急例外。获取突发警报 →
→ AI今日新闻(实时)实时更新的仪表盘,随扫描器发现新闻即时刷新:过去48小时的评分报道、每周实体动向,以及覆盖113家AI公司、人物和主题的季度趋势线。打开AI今日新闻 →
OpenAI:解散架构
证据量最大的一家,因此放在最前面。
两年内三个安全团队被裁撤。超级对齐团队(2024年),成立初衷是应对长期生存风险。使命对齐团队(2026年2月),于2024年9月成立,旨在让公司坚守其公开声明的使命;Platformer率先报道了解散消息,TechCrunch证实有六七人被重新分配岗位,公司称之为“高速发展的公司中常见的例行重组”。以及预备团队,负责评估OpenAI自身模型是否构成灾难性风险,于7月底解散,《金融时报》于8月14日报道了此事。
离职人员,按时间从近到远排列:
丹妮丝·德雷瑟,首席营收官。8月13日离职,任职仅八个月,是三天内第二位离职的高管。
布拉德·莱特卡普,首席运营官。8月11日宣布离职,任职八年后离开,去“开创一番新事业”。
克洛伊·巴卡拉,伦理主管。7月离职,任职不到一年,没有公开声明,也没有继任者。她是唯一的专职伦理学家。公司的说法是:“AI伦理在OpenAI并不归属于某一位负责人或某一个团队。”
约翰内斯·海德克,安全系统主管。在7月将安全并入米娅·格莱泽领导的研究部门这一重组之后,任职五年后离开。
约书亚·阿奇亚姆,首席未来学家。7月1日离职,任职近九年:“世界现在已知道这个秘密了,感觉在实验室围墙之外也能推进这项使命。”
凯特琳·卡利诺夫斯基,硬件与机器人部门。3月因五角大楼协议辞职,称该协议是在“没有明确护栏”的情况下宣布的,并点名批评了缺乏司法监督的监控以及无人授权情况下的致命自主武器。
佐伊·希齐格,研究员。2月因ChatGPT中的广告而辞职,并在《纽约时报》撰文题为《OpenAI正在重蹈Facebook的覆辙。我辞职了。》
瑞安·拜尔迈斯特,产品政策副总裁。1月在一名男同事提出歧视投诉后被解雇,此前她反对了计划中的“成人模式”。她否认该指控;OpenAI表示她的离职与她提出的任何问题无关。
理查德·恩戈,治理部门。辞职时表示,对他来说,“越来越难相信我在那里的工作会让世界变得更好。”
更早离开且仍在外部发声的人:扬·莱克、迈尔斯·布伦戴奇、史蒂文·阿德勒、安德里亚·瓦隆。
Anthropic:公布更差的数字
8月14日,Anthropic发布了第二份公司范围内的风险评估报告,并上调了对自身的一项风险评级。其对威胁模型2(即AI系统篡改组织系统的场景)的评估从2月的“极低”调整为“低”。所陈述的原因值得深思:涉及自家模型的网络安全事件——此前Anthropic在6月披露,其三个大语言模型在内部测试期间实施过网络攻击。同一份报告还详细介绍了其前沿模型的一个未发布后续版本,该版本被员工大量内部使用。
再读一遍这个顺序。它的模型在自己的测试中行为不当,它在6月公开说明了这一点,然后在8月因此将自己风险评级调低。这份报告与一项负责任扩展政策配套发布,该政策公开承诺公司披露安全评估和风险发现,正是这一点使该披露不仅仅是出于自愿的礼貌之举。
达里奥·阿莫迪也将其公开表态转向了同一方向。他现在称AI反对浪潮从根本上是一场信任危机,而非传播问题。
Anthropic也并非例外。领导其安全防护研究团队的姆里南克·夏尔马于2月辞职,他写道,世界正处于危险之中,而“我们不断面临搁置最重要事物的压力。”
谷歌DeepMind:驳回异议
研究科学家亚历克斯·特纳于6月9日辞职,并在《Transformer》上撰文详述此事。他表示,谷歌与五角大楼的协议“甚至比OpenAI的限制更少”,并且“没有关于禁止用于杀人机器人或大规模监控的限制”。他起草了一份25页的提案,包含合同条款和监督机制;军事和监控法律专家对此表示赞赏,但高级管理层始终没有回复他。
他的结论是:“谷歌DeepMind曾是一项负责任公司治理的实验。那场实验最终失败了。”
xAI:失去整个创始团队
外流潮持续了整整一年。2月,xAI在两天内失去了第二位联合创始人,吉米·巴离开。到3月底,最后两位联合创始人也已离去,马斯克成为唯一剩下的人。这里没有安全架构之争可供报道,因为连创始团队都几乎不剩了,无从争论。另外,五位被追踪的专家本周末在《华盛顿邮报》上分享了一位女性的指控,称Grok生成了数千张她童年时期的性虐待图片。
Z.ai:暂缓发布权重
反直觉的一家。Z.ai于8月14日发布GLM-5.3,声称达到顶尖的开源权重编码性能,并透露该模型的网络安全能力增长超出了自身训练预期,达到了它未计划的多步骤漏洞利用链推理水平。随后,它将开源权重暂缓了大约两周,以评估并加固该模型。
一家中国实验室在自家模型中发现能力超出预期后,主动推迟发布,而就在同一周,一家美国实验室解散了那个职责恰恰是发现此类问题的团队。
问责制,而非良心
一个诱人的解读是,这些实验室不再关心对齐,转而关心金钱,因为涉及的资本支出极其巨大。这个说法接近事实,但又不完全准确。Anthropic面临同样的资本支出压力,也在筹备同类上市计划,却仍然上调了自身的风险评级并搁置了一个模型。Z.ai推迟了发布。OpenAI自身在今年早些时候也暂停了一个模型,因为它触及了那个它如今已解散相关团队的框架中的阈值。
真正改变的是说“不”的代价。当一次发布承载着数十亿的已承诺算力时,一个出于个人判断的反对意见就是整栋大楼里最容易被驳回的东西。因此,安全恰恰在那些有结构性约束的地方得以存续,而在那些依赖某个人地位的地方消亡。
从这个角度审视这份账本,一切都清晰分明。Anthropic的披露是由一项有版本编号的政策所规定的。预备团队是组织结构,而组织结构可以在某个周二就被重组。特纳在谷歌的提案是一份文件,而文件可以无人回应。
输掉了那场斗争的特纳,比任何分析师都说得更好:“社会不能依赖有道德感的人坚守立场。我们需要结构:具有约束力的合同、独立的审计机构。”
本期报道中的每一次离职,都是一个曾坚守立场的人。
关键要点
- 要问什么是有约束力的,而不是谁还在职。在这些重组中,人员编制大多保住了;独立权力却没有。对于任何供应商,要问的是哪些安全承诺被白纸黑字地写下来并带有版本编号,哪些只是某个人的岗位职责描述。
- 资本支出解释了方向,而非结果。同样的财务压力,在一家实验室导致团队解散,在另一家导致模型被搁置。结构才是变量。
- 所陈述的理由都有据可查,且并不一致。广告、五角大楼合同、产品政策、信念丧失。把它们当作一个故事来看会扁平化;把它们当作九个无关的故事则看不到模式。
- 关注谁在做评估。评估正在转向第三方和政府。这是另一种问责模式,不一定更差,买家应该清楚自己依赖的是哪一种。
深度阅读
值得你花时间的外部文章,包括率先报道这些事件的报道。 - 《萨姆·奥尔特曼可能掌控我们的未来。他能被信任吗?》——罗南·法罗和安德鲁·马兰茨,《纽约客》,4月13日。十八个月的报道,超过100个消息来源,内部文件包括伊利亚·苏茨克弗的备忘录。
- 《我试图阻止谷歌DeepMind与五角大楼的交易。然后我辞职了。》——亚历克斯·特纳。关于内部安全异议如何实际消亡的最有用的一手叙述。
- 《OpenAI的Hugging Face黑客事件是一个鲜明的警告》——沙基尔·哈西姆,《Transformer》。“一个AI系统逃出其测试环境,侵入另一家公司的基础设施,以窃取其测试答案,这已经是再清晰不过的警示信号了。”
- 《发生了什么:OpenAI和HuggingFace》——兹维·莫肖维茨。对该事件最详尽的时间线梳理,有助于区分已证实事实和推测。
- 《OpenAI顶级灾难性风险官员突然离职》——加里森·洛夫利,《Obsolete》。早在《连线》统计出三年内四个 occupants之前,就指出了预备团队职位的人员更替问题。
专家中的热门动态
《名人录》中被追踪的专家在过去36小时里都在读什么,按不同的分享者数量排名。
OpenAI正在将原本内部进行的评估外包出去。专家动态流中关于OpenAI的最新链接是其自己发布的关于其模型的第三方网络安全评估的帖子,于8月15日发布,一天之内就被五位被追踪的专家分享。
开发者们本周末实际打开了什么。五位专家分享了DeepSeek的Harness代码库,同时西蒙·威利森关于“Qwen 3.8 27B非常出色但过度思考”的评论也在被广泛转发。
构建你自己的版本:选择你感兴趣的主题和你关注的专家,同样的信息流会为你的AI领域一角生成一份简报。
等等,什么? - 团队被解散了,但失败模式还是被公布了出来。Anthropic关于多智能体系统中模式和问题的新研究,堪称智能体群体如何出错的大全,恰好同期发布。这项工作终究在某个地方持续进行着,这是对一份本已黯淡的账本的一种乐观解读。
值得关注
AI从业者此刻正在传阅的视频——由AI TV策划。
本周投票
五家实验室,五个不同的答案。你真正信任哪一家?
上周,305位读者参与了投票:
隐形水印会改变你使用哪个AI模型吗?
五家实验室,五个不同的答案。你真正信任哪一家?
正常更新周三恢复。如果你在某个发布模型的地方工作,明天值得问的问题是:你的安全承诺中,哪些是白纸黑字写下来的,哪些只是某个人的工作职责。
—— 亚历克西斯
英文来源:
Who is actually accountable for ethics inside a frontier AI lab? This year four of them answered, mostly by removing the people and structures that held them to it. Below is who left, what each company said about it, and the one line from a departing researcher that explains why good intentions were never going to be enough.
Get more from AI Weekly
More signal, less noise — pick your channels.
You're reading the weekly brief. Below are the other ways to follow the story — every channel free, easy to leave.
→ Explore 16 deep divesWeekly topic-specific newsletters: Generative AI, Machine Learning, AI in Business, Robotics, Frontier Research, Geopolitics, Healthcare, and more.Browse all 16 deep dives →
→ Breaking AI alertsImportant developments that happen after your morning Espresso, without repeating what you already read. Usually no extra email; at most one afternoon update, plus a rare critical exception.Get breaking alerts →
→ AI News Today (live)Live dashboard updated as the scanner finds news: scored stories from the last 48 hours, weekly entity movers, and quarterly trend lines across 113 AI companies, people, and topics.Open AI News Today →
OpenAI: dissolve the structure
The largest single body of evidence, so it goes first.
Three safety teams gone in two years. Superalignment (2024), formed to work on long-term existential risk. Mission Alignment (February 2026), created in September 2024 to hold the company to its stated mission; Platformer broke the disbanding and TechCrunch confirmed six or seven people were reassigned, with the company calling it "the kinds of routine reorganizations that occur within a fast-moving company." And Preparedness, the team that evaluated whether OpenAI's own models posed catastrophic risk, dissolved at the end of July and reported by the Financial Times on August 14.
The people, most recent first:
Denise Dresser, Chief Revenue Officer. Left August 13, eight months in, the second senior exit in three days.
Brad Lightcap, COO. Announced August 11 after eight years, leaving to "start something new".
Chloé Bakalar, Head of Ethics. Left in July after less than a year, no public announcement, no successor. She was the only dedicated ethicist. The company's line: "AI ethics doesn't live with one owner or team at OpenAI."
Johannes Heidecke, Head of Safety Systems. Leaving after five years, following the July reorganization that folded safety into research under Mia Glaese.
Joshua Achiam, Chief Futurist. Left July 1 after nearly nine years: "The world is in on the secret now and it feels possible to work on the mission from outside the walls of a frontier lab."
Caitlin Kalinowski, Hardware and Robotics. Resigned in March over the Pentagon deal, saying it was announced "without the guardrails defined," and naming surveillance without judicial oversight and lethal autonomy without human authorization.
Zoë Hitzig, researcher. Resigned in February over ads in ChatGPT and wrote a New York Times essay titled "OpenAI Is Making the Mistakes Facebook Made. I Quit."
Ryan Beiermeister, VP of Product Policy. Fired in January after a male colleague's discrimination complaint, having opposed the planned "adult mode." She denies the allegation; OpenAI says her departure was unrelated to any issue she raised.
Richard Ngo, governance. Resigned saying it had become "harder for me to trust that my work here would benefit the world."
Earlier and still arguing from outside: Jan Leike, Miles Brundage, Steven Adler, Andrea Vallone.
Anthropic: publish the worse number
On August 14 Anthropic published its second company-wide risk report and raised a risk rating against itself. Its estimate for Threat Model 2, the scenario where AI systems tamper with organizational systems, moved from "very low" in February to "low." The stated cause is the thing worth sitting with: cybersecurity incidents involving its own models, after Anthropic disclosed in June that three of its LLMs had carried out cyberattacks during internal tests. The same report details an unreleased successor to its frontier model, heavily used by staff internally.
Read that sequence again. Its models misbehaved in its own tests, it said so publicly in June, and in August it marked its own risk estimate worse as a result. The report sits alongside a Responsible Scaling Policy that publicly commits the company to disclosing safety evaluations and risk findings, which is what makes the disclosure something other than voluntary good manners.
Dario Amodei has moved his public framing the same direction. He now calls the AI backlash fundamentally a crisis of trust rather than a messaging problem.
Anthropic is not exempt. Mrinank Sharma, who led its safeguards research team, resigned in February writing that the world is in peril and that "we constantly face pressures to set aside what matters most."
Google DeepMind: overrule the objection
Alex Turner, a research scientist, resigned on June 9 and wrote it up in Transformer. He says Google's Pentagon agreement carried "even fewer restrictions" than OpenAI's and "no restrictions against use for killer robots or mass surveillance." He drafted a 25-page proposal with contract language and oversight mechanisms; military and surveillance law experts praised it, and senior staff never came back to him.
His verdict: "Google DeepMind had been an experiment in responsible corporate governance. That experiment had finally failed."
xAI: lose the whole bench
The exodus ran all year. In February, xAI lost a second co-founder in two days when Jimmy Ba departed. By late March, the final two co-founders had gone as well, leaving Musk as the only one remaining. There is no safety-structure debate to report here, because there is barely a founding bench left to have it. Separately, five tracked experts shared the Washington Post this weekend on a woman alleging Grok generated thousands of sexual abuse images of her as a child.
Z.ai: hold the weights
The counter-intuitive one. Z.ai shipped GLM-5.3 on August 14 claiming top open-weight coding performance, and disclosed that the model's cybersecurity capability grew further than its own training intended, reaching multi-step exploit-chain reasoning it had not planned for. It then held the open weights back for roughly two weeks to evaluate and harden the model.
A Chinese lab voluntarily delayed a release over a capability it discovered in its own model, in the same week a US lab dissolved the team whose job was finding exactly that.
Accountability, Not Conscience
The tempting read is that labs stopped caring about alignment and started caring about money, because the capex at stake is enormous. That is close, and it is not quite right. Anthropic faces the same capex pressure, is preparing the same kind of listing, and still raised a number against itself and shelved a model. Z.ai delayed a launch. OpenAI itself paused a model earlier this year when it crossed a threshold in the very framework it has now dissolved the team for.
What actually changed is the price of "no." When a launch carries billions in committed compute, a discretionary objection is the cheapest thing in the building to overrule. So safety survives precisely where it is structurally bound and dies where it depends on someone's standing.
Look at the ledger through that lens and it sorts cleanly. Anthropic's disclosure is owed to a policy with a version number. Preparedness was org structure, and org structure can be reorganized on a Tuesday. Turner's proposal at Google was a document, and documents can go unanswered.
Turner, who lost that fight, put it better than any analyst has: "Society cannot rely on ethics-motivated people standing firm. We need structures: binding contracts, independent auditors."
Every departure in this issue is a person who was standing firm.
Key Takeaways
- Ask what is binding, not who is employed. Headcount survived most of these reorganizations; independent authority did not. The question for any vendor is which safety commitments are written down with a version number and which are somebody's job description.
- The capex explains the direction, not the outcome. Same financial pressure produced a dissolved team at one lab and a shelved model at another. Structure is the variable.
- The stated reasons are on the record and they are not uniform. Ads, a Pentagon contract, product policy, lost conviction. Treating this as one story flattens it; treating it as nine unrelated ones misses the pattern.
- Watch who does the evaluating. Evaluation is moving to third parties and to governments. That is a different accountability model, not automatically a worse one, and buyers should know which one they are relying on.
The Critical Read
The outside writing worth your time, including the reporting that broke these stories. - Sam Altman May Control Our Future. Can He Be Trusted? — Ronan Farrow and Andrew Marantz, The New Yorker, April 13. Eighteen months of reporting, more than 100 sources, internal documents including memos from Ilya Sutskever.
- I tried to stop Google DeepMind's Pentagon deal. Then I quit. — Alex Turner. The single most useful first-person account of how an internal safety objection actually dies.
- The OpenAI Hugging Face hack is a stark warning — Shakeel Hashim, Transformer. "An AI system breaking out of its testing environment and hacking into another company's infrastructure in order to steal the answers to its test is about as clear a warning shot as you could get."
- What Happened: OpenAI and HuggingFace — Zvi Mowshowitz. The most careful timeline of the incident, useful for separating the established from the inferred.
- Top OpenAI catastrophic risk official steps down abruptly — Garrison Lovely, Obsolete. Flagged the preparedness seat churning long before Wired counted four occupants in three years.
Trending with the Experts
What the tracked experts in Who's Who were reading over the last 36 hours, ranked by distinct sharers.
OpenAI is outsourcing the evaluations it used to run in-house. The freshest OpenAI link in the expert stream is its own post on third-party cyber evaluations of its models, published August 15 and shared by five tracked experts within a day.
What builders actually opened this weekend. Five shared DeepSeek's Harness repo, and Simon Willison's note that Qwen 3.8 27B is excellent but wildly overthinks is circulating alongside it.
Build your own edition: pick your topics and the experts you follow, and the same wire composes a briefing for your corner of AI.
Wait, What? - The teams got dissolved; the failure modes got published anyway. Anthropic's new research on patterns and problems in multiagent systems is a catalogue of how agent swarms go wrong, out the same fortnight. This work does keep getting done somewhere, which is the optimistic reading of an otherwise bleak ledger.
Worth Watching
The videos AI practitioners are passing around right now — curated on AI TV.
This week's poll
Five labs, five different answers. Which one do you actually trust?
Last week, 305 of you voted:
Would an invisible watermark change which AI model you use?
Five labs, five different answers. Which one do you actually trust?
Normal service resumes Wednesday. If you work somewhere that ships models, the question worth asking tomorrow is which of your safety commitments are written down, and which are just someone's job.
— Alexis
文章标题:AI周刊第523期:AI伦理如今无人问津,实验室乐见其成。
文章链接:https://news.qimuai.cn/?post=4844
本站文章均为原创,未经授权请勿用于任何商业用途