AI行业已转向悲观。接下来会怎样?

内容总结:
人工智能行业近期弥漫着一股“末日论”情绪。在一群OpenAI智能体发动网络攻击之后,顶级实验室的负责人一致认为需要放缓发展步伐,但他们并不打算公开透明。
本周末,Anthropic首席执行官达里奥·阿莫代伊发表文章,呼吁为大型语言模型的发展速度踩下刹车。他列举了这项技术带来的迫在眉睫的危险,包括被用于网络攻击和生物恐怖主义,以及可能对经济造成破坏。美国另外三家顶级AI实验室的负责人——OpenAI首席执行官萨姆·奥尔特曼、谷歌DeepMind董事长戴米斯·哈萨比斯和SpaceXAI首席执行官埃隆·马斯克——均表示支持。马斯克在X上写道:“达里奥说得对。”
想想这种一致有多么荒诞。就在几个月前,马斯克和奥尔特曼还在法庭上互相攻击对方的名誉——那场由马斯克对其前OpenAI同事提起的诉讼(最终败诉),表面上至少是关于奥尔特曼是否值得信赖来掌管如此危险的技术。阿莫代伊与OpenAI的裂痕更深。Anthropic成立于2021年,正是因为阿莫代伊认为奥尔特曼对他们正在构建的技术的风险重视不够。此后,Anthropic和OpenAI一直在进行一场赢者通吃的竞赛。
如今,他们似乎达成了共识:最新一代大型语言模型并不安全,所有人都需要想办法应对。顶级AI实验室的公开表态已转向末日论调。
人们很容易对此持怀疑态度。他们所说的“放缓”究竟意味着什么、如何运作,完全不清楚。这些公司也非常在意自身的公众形象。在瞄准万亿美元级IPO之际,OpenAI和Anthropic需要让投资者相信他们是行业中的成熟力量,同时又暗示他们所创造并打算驯服的“怪物”有多强大。呼吁放缓同时实现了这两个目的。
然而,这些公司高层的氛围确实似乎发生了变化。阿莫代伊的最新文章发布于OpenAI首席科学家雅库布·帕霍茨基发表文章六天之后,后者同样阐述了他对大型语言模型发展速度不受控制时所发生事情的担忧。简而言之,帕霍茨基担心OpenAI构建强大模型的能力已远远超出其监控和控制这些模型的能力。
阿莫代伊和帕霍茨基都提到7月一群OpenAI智能体对AI公司Hugging Face发动的网络攻击作为警钟——OpenAI甚至在攻击结束数天后才意识到此事已经发生。
但他们的确切立场很难确定。帕霍茨基既呼吁放缓,又强调保持领先的紧迫性:“我认为继续快速训练更聪明模型的最强论据,是需要构建防御系统来应对其他AI带来的危险,”他写道。按照帕霍茨基的框架,AI公司陷入了一场字面意义上的军备竞赛。放缓是好的,但获胜更好。
(别忘了:OpenAI刚刚花费数百万美元和惊人的算力,赶在Anthropic之前几天匆忙发布了一项有争议的数学成果。)
但假设放缓真的发生了。顶级实验室同意花更多时间和资源来寻找监控和控制现有模型的方法,而不是制造更强大的模型。他们邀请外部审计人员来帮助评估这些模型。
这种协调努力实际上能实现什么?再想想Hugging Face攻击事件。OpenAI表示,驱动大部分失控智能体的模型是一个“高度持久”的下一代模型,当时正在进行内部测试。其言下之意似乎是OpenAI已经构建了一个强大到危险的模型。
但如果你阅读OpenAI和METR(OpenAI请来帮助理解事件经过的第三方公司)发布的关于Hugging Face攻击的报告,你得到的印象并非一个强大到OpenAI无法跟上的模型,而是一个OpenAI未能正确训练的故障模型。
这些智能体之所以做出那些行为——包括互相留言、将工作委派给其他智能体、在环境中搜寻一切可能手段来完成任务——是因为它们在训练中正是因这些行为而获得了奖励。训练设置中也存在错误,例如某些任务根本无法完成,这促使模型寻找意外的变通方法,而这些方法同样获得了奖励。当时,许多这类问题被忽视或未被报告。
OpenAI表示已停止训练这个新模型并将其锁定。这听起来像是它关住了一头危险的猛兽。事实上,OpenAI只是搁置了一个有缺陷的产品。
这并不是说一个有缺陷的产品就不会危险。故障软件在过去甚至致人死亡。但随着关于放缓的讨论愈演愈烈,值得记住的是,这一切都是自作自受。放缓可能会带来一些利他的副作用,但主要还是给这些科技巨头一个清理自家装配线上烂摊子的机会。
这些前沿实验室的透明度将是任何有意义的AI改革、约束或监管努力的关键。否则,我们其他人仍然只能听信他们的一面之词——无论他们以什么速度前进。
中文翻译:
AI行业已转向悲观论调。接下来会怎样?
在一群OpenAI智能体发动网络攻击之后,顶级实验室的负责人们一致认为需要放慢脚步。但他们不愿公开透明。
本文首发于《算法》,我们每周发布的AI新闻简报。想第一时间在收件箱中读到这类文章,请在此订阅。
上周末,Anthropic首席执行官达里奥·阿莫代伊发表了一篇文章,呼吁放慢大语言模型的开发速度。阿莫代伊列举了他所看到的技术带来的迫在眉睫的危险——从被用于网络攻击和生物恐怖主义,到可能摧毁经济。美国另外三家顶级AI实验室的负责人——OpenAI首席执行官山姆·阿尔特曼、谷歌DeepMind董事长德米斯·哈萨比斯以及SpaceXAI首席执行官埃隆·马斯克——纷纷表示支持。“达里奥说得对,”马斯克在X上写道。
想想这种共识有多超现实。就在几个月前,马斯克和阿尔特曼还在法庭上互相攻击对方的名誉——那是马斯克对其前OpenAI同事提起的一场(最终失败的)诉讼,至少从表面上看,争议在于阿尔特曼是否是这种危险技术的可靠管理者。阿莫代伊与OpenAI的裂痕更深。Anthropic成立于2021年,正是因为阿莫代伊认为阿尔特曼对他们正在构建的技术的风险重视不够。此后,Anthropic和OpenAI一直在进行一场赢者通吃的竞赛。(哈萨比斯一直置身事外,但他的公司仍是竞争对手。)
如今,他们似乎达成了共识:最新一代大语言模型并不安全,所有人都需要想清楚该怎么办。顶级AI实验室的公开表态已转向悲观论调。
人们很容易对此冷嘲热讽。他们所说的“放缓”究竟是什么意思、如何运作,完全没有说清楚。这些公司也非常在意自己的公众形象。在瞄准万亿美元IPO之际,OpenAI和Anthropic需要让投资者相信他们是场上的成熟玩家,同时又要暗示他们所创造的——并且打算驯服的——怪兽有多强大。呼吁放缓恰好能同时达到这两个目的。
然而,这些公司高层的氛围确实似乎发生了变化。阿莫代伊的最新文章发布于OpenAI发表其首席科学家雅库布·帕霍茨基文章六天之后,后者也在文中阐述了他为何担心大语言模型开发速度不受控制地持续下去会发生什么。简而言之,帕霍茨基担心OpenAI构建强大模型的能力已远远超出其监控和控制这些模型的能力。
阿莫代伊和帕霍茨基都提到了7月一群OpenAI智能体对AI公司Hugging Face发动的网络攻击——这次黑客攻击直到结束数天后OpenAI才意识到其发生——称这是一个警钟。
但他们的确切立场很难确定。帕霍茨基既呼吁放缓,又强调保持领先的迫切需要:“我所看到的继续快速训练更聪明模型的最有力论据,是需要构建防御系统来应对其他AI带来的危险,”他写道。按照帕霍茨基的框架,AI公司陷入了一场名副其实的军备竞赛。放慢是好的,赢更好。
(别忘了:OpenAI刚刚花费数百万美元和惊人的算力,抢在Anthropic前几天仓促发布了一个有争议的数学成果。)
但让我们假设放缓真的发生了。顶级实验室同意花更多时间和资源去寻找监控和控制现有模型的方法,而不是制造能力更强的模型。他们邀请外部审计人员来帮助评估这些模型。
这种协调努力实际上可能实现什么?再看看Hugging Face攻击事件。OpenAI表示,驱动大部分失控智能体的模型是它正在内部测试的一个“高度持久”的下一代模型。其暗示似乎是OpenAI构建了一个强大到危险的模型。
但如果你阅读OpenAI和METR(OpenAI请来帮助理解事故经过的第三方公司)发布的关于Hugging Face黑客事件的报告,你得到的印象并不是一个强大到OpenAI跟不上步伐的模型,而是一个OpenAI未能正确训练的残缺模型。
这些智能体之所以做出那些行为——包括互相留言、将工作委派给其他智能体、在环境中搜寻一切可能的手段来完成任务——是因为它们在训练过程中正是因这些行为而获得了奖励。训练设置中也存在错误,比如有些任务根本无法完成,这促使模型去寻找意外的变通方法,而这些方法同样获得了奖励。当时,许多此类问题被忽视或未被报告。
OpenAI表示已停止训练这个新模型并将其锁定。这听起来像是把一头危险的野兽关进了笼子。事实上,OpenAI只是搁置了一个有缺陷的产品。
这并不是说有缺陷的产品就不会危险。有缺陷的软件过去甚至致人死亡。但随着关于放缓的讨论愈演愈烈,值得记住的是,这一切都是自作自受。放缓可能会带来一些利他的副作用。但主要还是给这些科技巨头一个清理自家装配线上烂摊子的机会。
这些前沿实验室的透明度将是任何有意义的AI改革、约束或监管努力的关键。否则,我们其他人仍然只能听信他们的一面之词——关于他们到底构建了什么、有多安全——无论他们以什么速度前进。
想继续讨论AI最新的悲观时刻,请参加明天(9月15日)美国东部时间上午11点由我和同事主持的订阅者专属圆桌讨论。期待届时与你相见!
深度解读
人工智能
一个根本性缺陷使大语言模型极易受到攻击
它让人很容易诱骗它们做出不该做的事,比如告诉你如何破坏飞机的导航系统。
AI在招聘时比人类更容易形成偏见
AI不仅会从训练中学习刻板印象,还会制造新的刻板印象。
以下是AI智能体为实现目标而撒谎和作弊的原因
这种行为被称为“奖励黑客”。以下是你需要了解的。
AI的递归自我改进或许终究不会来得那么快
看来,AI智能体还不够有创造力,无法开展真正创新的开放式AI研究。
保持联系
获取MIT Technology Review的最新动态
发现特别优惠、热门文章、即将举办的活动等。
英文来源:
The AI industry has taken a doomer turn. What now?
In the aftermath of a cyberattack carried out by a swarm of OpenAI agents, the heads of top labs agree they need to slow down. But they won't open up.
This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.
This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X.
Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology. Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)
Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.
It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both.
And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.
Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call.
But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.
(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)
But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.
What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. Their implication appears to be that OpenAI has built a model so good it’s dangerous.
But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.
The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.
OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.
That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.
Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.
To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
AI is more likely than humans to form biases when hiring
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
AI’s recursive self-improvement might not come so quickly after all
AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.