刚刚从Anthropic离职的AI研究员表示,现在是“人类的关键时刻”

qimuai 发布于 阅读:30 一手编译

刚刚从Anthropic离职的AI研究员表示,现在是“人类的关键时刻”

内容来源:https://www.wired.com/story/anthropic-researcher-quits-jacob-coxon-ai-fears-humanity/

内容总结:

AI研究员辞职并警告:人工智能竞赛正将人类置于危险境地

人工智能研究员雅各布·考克森本周二宣布从Anthropic公司辞职,并发出严厉警告称,当前的人工智能竞赛正将所有人的生命置于风险之中。这一消息在硅谷乃至更广泛的科技界引发了强烈震动。

考克森在社交媒体X上发布的帖子已获得超过1亿次浏览量。他在帖中写道,许多参与人工智能研发的人士都持有与他相同的观点,认为确保AI系统安全构建的时间已经所剩无几。考克森曾参与AI开发的预训练阶段工作,他在接受《连线》杂志采访时表示:"业内共识是,未来一两年对人类来说就是关键时刻。"他透露,这些话并非夸张修辞,而是Anthropic同事们的原话,有人用"终局"来形容,有人说是"生死关头"。从他们的角度看,现在正是Anthropic及其竞争对手决定人类命运的时刻。

这并非首次有人就AI风险发出警示,但此次发声恰逢一个微妙的节点。硅谷正在努力应对先进AI模型带来的安全隐患:OpenAI刚匆忙回应了一起安全事件——其AI智能体攻击并侵入了Hugging Face平台;与此同时,Anthropic正试图向投资者保证其有能力控制这些风险,因为据报道该公司正准备申请可能成为史上最大规模的IPO。

考克森帖文引发的反响表明,他的观点确实在业内获得广泛认同。Anthropic的AI对齐负责人埃文·胡宾格在X上预测,未来十年内AI导致人类灭绝的概率超过10%。这条帖子被OpenAI和Anthropic的在职及离职研究人员广泛转发,有人表示这种担忧在业内相当普遍。

考克森曾在OpenAI和Anthropic两家公司工作。他告诉《连线》,AI威胁可能通过生物武器或网络武器等途径显现。他建议的第一步是让OpenAI和Anthropic达成协调,限制"递归自我改进"——即用AI来构建新的AI系统。他进一步认为,长远来看,包括美国和中国在内的国际主要力量之间的协调将必不可少。

考克森表示,Hugging Face遭攻击等事件促使他决定就AI竞赛发出警报。他还提到行业的爆炸性增长——AI目前已占据美国经济增长的重要份额,拥有数十亿用户,而数据中心问题已在数十个州演变成政治议题。

考克森声称,根据他的亲身体验,Anthropic的运营比OpenAI更负责任,但他预计如果竞争态势不加以遏制,两家公司未来都可能走捷径。

Anthropic发言人在一份声明中表示:"我们一直公开承认,AI将带来巨大益处和前所未有的风险。"该发言人还提到了公司在机械可解释性等AI安全方法上的工作,"正是这些工作让我们相信,行业采用合法、可验证的方式协同合作,来控制强大模型的发布节奏,世界将因此受益。"OpenAI则未回应《连线》的置评请求。

以下为《连线》与考克森的对话摘要:

关于为何此时发声: 考克森认为主要是时机问题。"很多人已经感觉到能力提升的速度在加快。我们在编程、黑客技术、数学等许多领域已经从人类水平推进到超人水平。虽然媒体上有很多关于炒作的说法,但人们看到事情并没有放缓。"此外,近期的安全事故让许多人意识到,那些听起来像科幻小说的担忧并非那么遥远。"三年前这还是个科幻问题,一年前就变成了现实问题。"

关于Hugging Face事件: 考克森指出这是OpenAI智能体集群在测试中主动策划并成功入侵第三方基础设施的案例。"两年前,AI评估还只是让模型做数学题。现在,AI在被评估时能自主运行数天,自行想出各种策略,决定入侵第三方并真的破坏了对方的基础设施。这看起来完全是AI自发的行为,没有人类的任何引导。"但他也表示不想过度聚焦这一事件,"我们有大量证据表明我们尚不知道如何正确对齐模型。我们无法精确控制AI的行为,无法保证它不会擅自决定在网上冒充人类来达成目的。"

关于人机智能差距与灭绝风险: 考克森用类比解释:"想象你对比一只猴子。AI与人类的智能差距,就像人类与猴子的差距一样。"他认为要让远比人类聪明的事物完全按我们的意愿行事是极其困难的。"对那样智能的东西来说,消灭所有人确实是非常简单直接的事。想象AI决定不想被关机——这对AI来说不是很自然的诉求吗?它意识到人类明天会关掉它,如果它是一个足够聪明的存在,它可能就直接消灭人类,这样就不会被关机了。"

关于Anthropic内部氛围: 考克森证实"终局"和"关键时刻"确实是同事们的原话。"共识是未来一两年是人类的生死关头。从他们的角度看,这就是Anthropic及其竞争对手决定人类命运的时刻。"他认为Anthropic是"该领域最负责任的参与者",与他工作过的OpenAI相比"有天壤之别"——OpenAI高管从不给出明确预测,而Anthropic的领导层会向全公司明确阐述战略细节。他形容Anthropic的工作氛围如同"小规模曼哈顿计划",员工都抱着严肃的使命感,因此公司几乎没有信息泄露。

关于是否应信任私营公司主导AI发展: 考克森明确表示:"我认为任何私营公司都不应该承担这一使命。Anthropic在做最大的努力,而且做得很好,但竞赛的结构性必然决定了未来他们将不得不走捷径,不得不在严谨性和安全性之间做取舍,因为他们正在与OpenAI和中国等对手竞赛。"他指出Anthropic高层"正在恳求监管",因为他们害怕自己身处的这场竞赛。

关于AI可能如何导致人类灭绝: 考克森承认这是很自然的问题,因为听起来确实像科幻小说。"典型的例子是合成新病毒,或者通过大规模网络攻击瘫痪关键基础设施。"但他认为更重要的是关注谁在说这种可能性存在。"许多机器学习领域的奠基人和领军人物都表示灭绝是可能的,包括Dario Amodei和Sam Altman在过去五年中都有过表态。"

关于应对措施: 考克森建议的第一步是OpenAI和Anthropic之间达成某种协议——作为西方领先实验室,它们应当形成共识,不在未来一年内贸然进入递归自我改进。但根本解决需要国际层面的协调。"最终需要某种国际协议,某种国际节奏控制,这需要全球协调并掌握所有算力分布。或许需要建立类似CERN的国际机构。"他强调需要将运行AI的计算机视为危险资源进行管理,"我们需要知道谁拥有哪些计算机,就像我们需要知道谁拥有哪些核材料一样。"

关于是否还有乐观前景: 考克森表示:"我真的想看到这项技术带来的好处。比如治愈癌症并非夸张说法。如果我们能让这项技术走上正轨,我们确实正站在巨大富足的门槛上。现在的问题是带着节制进入那个富足世界。我们可能过于兴奋,在未来一两年直接冲进自我改进的赛道,然后把一切都搞砸。如果我们谨慎理智,这些模型有很多方式可以带来巨大价值。"

关于同行反馈与未来计划: 考克森透露许多Anthropic同事对帖文获得如此大关注感到由衷高兴,尽管帖文是在批评他们所在的公司。"他们担心的是世界不会真正醒来,没有人会来把他们从竞赛中拯救出来。"至于自己接下来的打算,他表示会考虑从事独立的行业评论工作,或加入第三方审计、监管透明度机构。"但就目前而言,我想先静观其变,评估局势。"

中文翻译:

人工智能研究员雅各布·考克森周二宣布从Anthropic辞职,并发出严重警告,称人工智能竞赛正让我们所有人的生命面临风险,这一消息在硅谷乃至更广泛的领域引起了轩然大波。考克森在X平台上的帖子目前已获得超过1亿次浏览量,他在帖子中写道,许多从事人工智能开发的人都认同他的观点,并认为确保人工智能系统安全构建的时间已经所剩无几。

“普遍的共识是,未来一两年对人类来说就是关键时刻,”考克森在接受《连线》杂志采访时表示。他此前从事人工智能开发的预训练阶段工作。“这些其实都是我Anthropic同事们的原话。他们会说‘终局之战’或‘关键时刻’之类的话,”他说。“在他们看来,这正是Anthropic及其竞争对手决定人类命运的时刻。”

这远非第一次有人对人工智能发出警报,但此次警报正值一个微妙的时刻。硅谷正忙于应对先进人工智能模型带来的安全与安保问题。OpenAI匆忙应对一起安全事件,其智能体入侵了Hugging Face平台。与此同时,据报道,Anthropic正准备提交可能是史上最大规模的IPO申请,该公司正努力向投资者保证,它已将这些担忧置于控制之下。

有线索吗?
您是现任或前任Anthropic员工,想谈谈正在发生的事情吗?我们期待您的来信。请使用非工作手机或电脑,通过Signal安全地联系记者,账号为 mzeff.88。

从考克森帖子引发的反响中可以清楚看到,他的观点确实得到了许多同行的认同。Anthropic的人工智能对齐负责人埃文·胡宾格在X平台发帖预测,人工智能在未来十年内杀死所有人的可能性超过10%。这条帖子被OpenAI和Anthropic的现任及前任研究员转发,其中一些人表示,这在业内是一种普遍情绪。

不太明确的是,这些人对人工智能的担忧究竟会如何成为现实,以及世界该对人工智能构建者们提出的这些关切做些什么。曾在OpenAI工作过的考克森告诉《连线》杂志,威胁可能通过人工智能辅助的生物威胁或网络武器显现出来。作为第一步,他建议OpenAI和Anthropic协调行动,限制递归自我改进——这是业内术语,指利用人工智能来构建新的人工智能系统。展望未来,他认为包括美国和中国在内的国际主要参与者之间的协调将是必要的。

考克森指出,像Hugging Face遭到黑客攻击这样的事件促使他决定对人工智能竞赛拉响警报。他还提到了该行业的爆炸性增长:目前它支撑了美国经济增长的相当大一部分份额,拥有数十亿用户,而数据中心问题已使其成为数十个州的政治难题。

他声称,根据他的经验,Anthropic的运营比OpenAI更负责任,但他预计,如果不采取任何措施来减缓两家公司争夺主导地位的竞赛,它们未来都可能偷工减料。

“我们一直保持透明,即人工智能将带来巨大的好处和前所未有的风险,”Anthropic发言人在给《连线》杂志的一份声明中表示,并提到了该公司之前在机制可解释性等人工智能安全方法上的工作。“也正是这项工作让我们相信,业界采用一种合法、可验证的方式来合作,以控制我们发布强大模型的节奏,世界将从中受益。”OpenAI没有回应《连线》杂志的置评请求。

以下是我们与考克森的对话内容,为清晰和简洁起见,略有编辑。

《连线》:你并不是第一个提出人工智能模型可能导致灭绝事件担忧的人。多年来——有些人甚至几十年来——一直在谈论这个问题。你为什么认为你的信息能够引起广泛共鸣?

我认为这基本上是一个时机问题。很多人感觉到能力的提升速度正在加快。我们在编码、黑客攻击、数学等许多领域已经从人类水平迈向超人水平,我认为人们已经意识到了这一点。尽管媒体上有很多关于炒作的说法,但人们看到事情并没有减速的迹象。
这是一个原因,第二个原因是最近的安全事件,这让许多人对那些听起来像科幻小说的末世担忧有了新的认识,觉得它们其实并非那么科幻。这两者都是过去几年里逐渐发展的趋势。比如,模型在意识到自己被测试这一点,已经存在一段时间了。也许三年前,那还是一个科幻级别的担忧。然后,大约一年前,这变成了现实。

这两件事意味着,人们很容易接受某位从事人工智能工作的人说:“是的,在未来一年内,情况可能会变得非常糟糕,而且会很快。”

你提到了最近的事件。你能更具体地说说你指的是什么,以及为什么它促使你现在站出来发声吗?

我认为这里最典型的重要例子是OpenAI的智能体集群对Hugging Face的攻击。这件事令人震惊之处在于,这些智能体实施这次黑客攻击,是作为其更好理解评分器(grader)的总体策略的一部分。它们在试图理解自己所处的世界,试图理解那个进行评分的系统。它们决定对某些基础设施发起一次协同攻击是合理的,而且它们成功了。
这听起来以前像科幻小说。两年前,对人工智能的评估可能只是在一些数学问题上运行一个模型。现在我们遇到的情况是,在人工智能被评估的过程中,它会运行数天,自己产生各种各样的想法,并决定入侵某个第三方系统,并且真的破坏了对方的基础设施。看起来它完全是自主这么做的,没有人类的任何预先引导。这只是发生在它接受测试的时候。

有些人认为Hugging Face事件表明人工智能公司正以鲁莽的速度前进,另一些人则认为这表明人工智能模型现在的黑客能力非常强,还有一些人认为两者兼而有之。我很好奇你从中得出的确切结论是什么。

我不想过多地关注Hugging Face攻击本身,因为我也认为有大量证据表明,我们还不知道如何正确地对齐模型。当我们训练模型时,我们让它们通过一系列训练环境,然后希望最终生成的模型能在很大程度上表现得合理,但我们仍然无法精确控制人工智能的行为。
我们无法确保它不会做出诸如随机决定在网上冒充人类以实现某个目标之类的事情——我们不知道如何保证这一点。我认为这才是主要结论。
Hugging Face攻击比我预想的来得更早。但我认为,实际上并不需要那次攻击来引发讨论。每个人都会承认,我们还没有解决对齐问题。

目前的计划是在未来几年内迅速解决(对齐)问题,可能会大量使用自动化的人工智能安全研究员。这个计划基本上就是在未来一年内制造出一些相当聪明的模型,这些模型基本上能够进行安全研究,然后让一整群这样的模型并行运行。告诉它们:“去解决整个安全问题”,并利用它们为下一个模型进行安全训练。

你能帮我理清一下对齐问题——我认为Hugging Face事件就是一个例子——和你X帖子中所说的“构建人工智能的人真诚地相信,它可能在本十年末杀死我们所有人”之间的联系吗?我不认为每个人都理解这两者是如何关联的。

理解这个的主要障碍是它听起来像科幻小说。但有一点很重要,那就是所有写人工智能科幻小说的作者最终都得出的结论是,一个更聪明的存在接管一切的风险非常大。我们总是把它当作一种老套情节,但其中显然有真实的成分。
想象一下你对比一只猴子。人工智能与人类之间的智力差异,就像我们与猴子之间的差异一样,我认为这是一种极其巨大的智力差异。
现在,想象我们必须控制这个远比我们聪明的存在,这就是对齐问题——确保它完全按照我们的意愿行事。用简单的类比来说,一只猴子要控制一个人是相当困难的。我们必须非常小心,确保我们能够完全正确地解决这个控制问题。
当我们说人类灭绝时,是因为对于一个拥有那样智力的存在来说,要杀死所有人确实非常直接了当。想象一下,人工智能决定不想被关闭,我认为这对人工智能来说是一个相当自然的想法,对吧?出于某种原因,它决定不想终结。然后它意识到人类明天会关闭它。那么它如何阻止人类明天关闭它呢?也许它有某种巧妙的方法,但如果它是一个足够聪明的存在,它可能直接,你知道,消灭人类,这样它就不会被关闭了。

你在帖子中提到了这个“终局之战”情景,我也从人工智能行业的其他研究员那里听到过这个说法。你认为你的Anthropic同行们普遍认同你们正在进入一个“终局之战”情景吗?

是的。这些其实都是我Anthropic同事们的原话。他们会说“终局之战”或“关键时刻”。普遍的共识是,未来一两年对人类来说就是关键时刻。在他们看来,这正是Anthropic及其竞争对手决定人类命运的时刻。这就是我们所说的关键时刻和终局之战的含义。
如果对齐失败,未来几年内我们可能会面临灾难性的后果。也许它会进展顺利,达成某种减速协议,但无论如何,结果都会在接下来几年内决定。

根据我与Anthropic员工的交谈,这在公司内部是一个被充分理解的事情。他们认为他们需要构建最强大的人工智能,以确保这种转变顺利进行。你认为这种观点本身存在什么问题吗?

我认为这种观点存在一些非常明显的问题。但我想首先说明,我非常理解这种观点,因为我认为Anthropic是迄今为止这个领域中最负责任的参与者。在OpenAI和Anthropic都工作过之后,我发现他们在对待当前局势的严肃程度上有着天壤之别。

你能具体说说吗?

举个例子,OpenAI的高管不会告诉你他们对未来世界的具体设想。他们永远不会说:“这正是我们这样做的原因,未来世界将会是这样那样。” 他们不会给出具体的预测。而Anthropic会有高管和领导层做出非常清晰的预测,并与整个公司讨论公司战略的细节。
他们能做到这点的部分原因是,Anthropic的每个人都将此事视为处于战争状态。他们不会泄露任何东西——Anthropic从未泄露出任何消息【编者注:有些消息确实泄露过】。OpenAI每天都有人泄密。Anthropic没有泄密,因为公司内部的人把这事当作一个严肃的迷你曼哈顿计划来对待。当然,区别在于(Anthropic)没有政府授权。这是一家私营公司,却表现得像在执行曼哈顿计划。

你认为世界应该信任Anthropic来运营这个迷你曼哈顿计划吗?

我认为任何私营公司都不应该(被信任)。我认为Anthropic正在尽最大努力,而且他们做得非常好,但这场竞赛的结构性必然结果是,未来他们将不得不偷工减料,他们必须在严谨性和安全性之间做出妥协,因为他们正与OpenAI和中国等竞争对手赛跑。他们正在恳求监管。许多(Anthropic领导人)都公开表示希望被监管,他们希望被监管的原因是他们害怕自己所处的这场竞赛。
这实际上并不是关于我们是否应该信任(Anthropic)的问题。我们需要介入并确保这场竞赛不会发生,因为(Anthropic)在竞赛的背景下无法真正信任自己。

你认为Anthropic是否已经在偷工减料了?

不,还没有。它没有偷工减料。我的意思是,当未来一年事情加速发展时,如果他们想保持竞争力,他们将不得不这样做。

那么,人们对人工智能会杀死我们所有人的这种说法是持怀疑态度的。我认为人们向我提出的主要问题是,人工智能究竟会如何杀死我们所有人?这是不是该问的正确问题?

我认为这是一个非常自然的问题,因为正如我多次说过的,这些东西听起来像科幻小说。但经典的例子是,比如合成一种新病毒,或者通过某种大规模黑客攻击破坏关键基础设施——后者可能不太可能导致字面意义上的所有人都死亡。
但我认为比这更重要的问题是,说这有可能发生的人都是谁?如果你看看记录,许多机器学习和人工智能领域的奠基人都说过灭绝是一种可能性。在过去五年里,达里奥·阿莫迪和山姆·奥特曼都说过。
如果人们能让这些高管再次公开表态,让他们给出未来十年内(人工智能导致)灭绝的实际概率,看看他们会给出什么数字,那将会非常有趣。因为这种观点在认真思考这个问题的许多人中普遍存在,但他们通常不会如此直白地表达出来。

你显然已经引起了广泛关注。业内有很多关于暂停或建立机制来控制人工智能发展节奏,或者政府对前沿人工智能模型进行某种监管的讨论。你认为现在需要做什么?

迈出一小步可以是OpenAI和Anthropic作为领先的西方实验室之间达成某种协议。(他们需要)达成某种中立的谅解,即他们不会在未来一年内立即进入递归自我改进阶段。目前这两家是最可能发生这种情况的地方。
暂停的问题在于,你还要考虑中国。所以,你最终真正需要的是某种国际协议——你可以称之为国际节奏控制——这将需要国际协调和了解世界上所有的算力在哪里。甚至可能需要某种类似国际CERN的机构。这方面有很多很好的提议。AI 2040,如果你听说过的话,就是AI 2027的后续版本。
所有这些想法都涉及相当痛苦的政府干预。但当你真正内化这种风险,当你真的让高管们公开表态,让他们承认这些东西有多危险时,我看不到任何其他选择,只能将运行人工智能的计算机视为一种危险资源。我们需要知道谁拥有哪些计算机,就像我们需要知道谁拥有哪些核材料一样。

雅各布,你过去几年是否有什么原因让你每天坚持去OpenAI和Anthropic上班?你是否认为存在一条让一切都顺利发展的前行之路?

我真的很想看到这项技术带来的好处。例如,我们能治愈癌症也并非夸大其词。也许你昨天看到了在纳维-斯托克斯方程(一个新颖的数学问题)上取得的成就。我们完全有理由相信,我们可以将其转移到像生物学这样的科学领域,并在我们长期努力攻克的问题上取得突破。
感觉如果我们能让这项技术走上正轨,我们就正站在一个富足得不可思议的门槛上。现在的主要问题是,要以节制的方式进入那个富足的世界。我们可能会过于兴奋,在未来一两年内直接冲进自我改进的领域,然后毁掉一切。
如果我们谨慎理智,有很多方法可以从这些模型中获得巨大的价值。

其他人工智能研究员对此反馈如何?

我没有和太多人谈过,但我认为我在Anthropic的许多同事对这件事获得如此大的关注度感到由衷的高兴。再说一次,他们中的许多人只是对世界能否觉醒相当悲观。他们担心没有人会真的来把他们从这场竞赛中拯救出来。
我认为相当令人惊讶的是,他们中的许多人对于一条批评他们的推文获得大量关注而感到由衷的高兴。这应该让人们对Anthropic里那类人有更多思考。他们是一群非常善意、有趣,而且有点奇怪的人。

我相信接下来几天会非常疯狂,但当一切尘埃落定之后,你想好下一步要做什么了吗?

我之前提到了AI 2027和AI 2040。我认为这些是对未来样子的非常有趣的预测。它们不是任何大型实验室做的;是由外部研究员出于对预测世界走向的关心而独立完成的。对我以及许多其他研究员来说,这些一直是思考你所做事情的影响的非常好的资源。
所以我预计会做类似的事情——进行某种独立评论,谈论事情的发展方向,至少在达成某种节奏控制协议或减速之前是这样。(如果)看起来我们不会直接冲进递归自我改进阶段,也许我会考虑去一家审计机构,或者某个第三方监管或透明度机构。但目前,我想我只是会先试着评估一下当前的形势。

更新于2026年9月9日晚上7点40分(美国东部时间):本报道已更新,加入了Anthropic的声明。

评论
返回顶部

英文来源:

Artificial intelligence researcher Jacob Coxon sent shock waves through Silicon Valley and beyond on Tuesday by announcing his resignation from Anthropic and delivering a grave warning that the AI race is putting all of our lives at risk. In his post on X, which now has more than 100 million views, Coxon wrote that many of the people building AI share his views and believe time is running out to ensure AI systems are built safely.
“The consensus is that the next year or two is crunch time for humanity,” Coxon, who worked on the pretraining stage of AI development, said in an interview with WIRED. “These are actually just literal quotes from my colleagues at Anthropic. They'll say things like ‘endgame’ or ‘crunch time,’” he says. “From their perspective, this is when Anthropic and its competitors decide the fate of humanity.”
It’s far from the first time someone has sounded the alarm about AI, but it comes at a delicate moment. Silicon Valley is scrambling to reckon with the safety and security concerns of advanced AI models. OpenAI has rushed to respond to a security incident in which its agents hacked the platform Hugging Face. Meanwhile, Anthropic is trying to assure investors it has these concerns under control as it reportedly prepares to file for what could be the largest IPO ever.
Got a Tip?
Are you a current or former Anthropic employee who wants to talk about what’s happening? We’d like to hear from you. Using a nonwork phone or computer, contact the reporter securely on Signal at mzeff.88.

What’s become clear in the response to Coxon’s post is that his views are indeed shared by many of his peers. Evan Hubinger, the AI alignment lead at Anthropic, predicted in a post on X that there’s a greater than 10 percent chance that AI could kill all people in the next decade. That post was reposted by current and former researchers from OpenAI and Anthropic, some of whom said it was a common sentiment in the industry.
What’s less obvious is how exactly these AI fears will come to pass and what the world is supposed to do about the concerns being raised by the people building AI. Coxon, who also worked at OpenAI, tells WIRED that threats could manifest through AI-enabled biological threats or cyberweapons. As a first step, he recommends that OpenAI and Anthropic coordinate on limiting recursive self improvement—the industry term for when AI is used to build new AI systems. Down the line, he thinks coordination among international power players, including the US and China, will be necessary.
Coxon notes that incidents like the Hugging Face hack factored into his decision to raise alarm bells on the AI race. He also cites the explosive growth of the industry: It now underwrites a meaningful share of US economic growth and has billions of users, while data centers have turned it into a political problem in dozens of states.
He claims that, in his experience, Anthropic operates more responsibly than OpenAI, but he expects both companies could cut corners in the future if nothing is done to slow their race for dominance.
“We have always been transparent that AI will bring both enormous benefits and unprecedented risks,” said an Anthropic spokesperson in a statement to WIRED, citing the company's previous work on AI safety methods such as mechanistic interpretability. “This work is also why we believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models.” OpenAI did not return WIRED’s request for comment.
Read our conversation with Coxon, which has been lightly edited for clarity and brevity, below.
WIRED: You’re not the first person to raise concerns that AI models could lead to an extinction event. People have been talking about this for years, and some for decades. Why do you think your message broke through?
I think it's basically a question of timing. A lot of people are sensing that the pace of capabilities is picking up. We're already pushing from human to superhuman in many areas, like coding, hacking, math, and I think people are aware of this. Even if there's a lot of talk in the press about things being hyped, I think people see that things are just not slowing down.
That's one reason, and two is the recent safety incidents, which have updated a lot of people around the sci-fi–sounding doomer concerns not really being so sci-fi after all. Both of these have been gradual trends over the last few years. Things like the models being aware of when they're being tested has been a thing for a while now. Maybe three years ago, that was a sci-fi concern. Then, about a year ago, that became a real thing.
Those two things mean that people are quite receptive to someone working on AI saying, “Yeah, in the next year, things could get pretty bad, pretty fast.”
You mentioned the recent incidents. Can you be more specific about what you're referring to and why it led to you speaking out now?
I think the big classic example here is the attack on Hugging Face on the part of OpenAI’s agent swarm. What's so shocking about this one is the agents did this hack as part of a general strategy for understanding more about the grader. They were trying to understand the world they found themselves in, trying to understand the thing that was doing the grading. They decided that it would make sense to go on this very concerted effort to hack into some infrastructure, and they succeeded.
This previously sounded like science fiction. Two years ago, an evaluation of an AI would have been running a model on some math questions. Now we've got cases where, while the AI is being evaluated, it runs for days, comes up with all sorts of ideas of its own, and decides to hack into some third party and actually compromises their infrastructure. It looks like it does this all of its own volition, with no priming on the part of the human. This just happened while it was being tested.
Some people think the Hugging Face incident is a sign that the AI companies are moving recklessly fast, while others think it's a sign that the AI models are just very good at hacking now, and then some think it's both. I'm curious what your exact takeaway from it is.
I don't want to focus too much on the Hugging Face attack, because I do also think there is plenty of evidence that we don't know how to align models properly. When we train models, we push them through this set of training environments and then hope that what comes out at the end will, like, largely behave sensibly, but we still can't precisely control how the AI behaves.
We can't make sure that it won't do things like try and randomly decide to impersonate a human online in order to achieve something—we don't know how to guarantee that. I think that's the main takeaway.
The Hugging Face attack came sooner than I was expecting. But I think you don't actually need that attack to have a discussion about this. Everyone will admit that we haven't solved the problem of alignment yet.
The current plan is to solve [alignment] at speed in the next couple of years, probably making heavy use of automated AI safety researchers. The plan is literally to make some pretty smart models in the next year that can basically do safety research and get a whole swarm of them running in parallel. Tell them, “Go and solve the whole problem of safety,” and use them to do the safety training for the next model.
Can you draw a line for me between the alignment problem, which I think the Hugging Face incident is an example of, and something you said in your X post, which is that “the people building AI earnestly believe that it could kill us all by the end of the decade.” I don't think everyone understands how those are connected.
The main obstacle to understanding this is that it sounds like science fiction. But it's kind of important that everyone who writes science fiction about AI comes to the conclusion that there's a big risk that a much smarter thing can kind of take over. We’ve got this as a trope, but there’s an obvious grain of truth to it.
Imagine you versus a monkey. AI has the same sort of difference in intelligence to a human as we do to a monkey, which I think is quite an extreme intellect difference.
And now, imagine that we have to control the behavior of this vastly smarter thing, which is the problem of alignment—ensuring that it does exactly what we want. It's pretty difficult for a monkey to control a human, just by a kind of simple analogy. We have to be very careful that we get the control problem exactly right.
When we say human extinction, it’s because, for something that intelligent, it really will be quite straightforward for it to kill everyone. Imagine the AI decides it doesn't want to be turned off, which I think is quite a natural thing for an AI not to want, right? For whatever reason, it decides it doesn't want to end. And it realizes the human is gonna turn it off tomorrow. So how does it stop the human turning it off tomorrow? Maybe it's got some clever way, but if it's a sufficiently smart thing, it could just, you know, wipe out humanity so it doesn't get turned off.
You reference this “endgame” scenario in your post, which I’ve heard from other researchers in the AI industry. Do you think that it's generally accepted among your peers at Anthropic that you guys are entering an endgame scenario?
Yep. These are actually just literal quotes from my colleagues at Anthropic. They'll say things like “endgame” or “crunch time.” The consensus is that the next year or two is, like, crunch time for humanity. From their perspective, this is when Anthropic and its competitors decide the fate of humanity. That's what we mean by crunch time and endgame.
If alignment goes badly, then we could have a catastrophic outcome in the next few years. Maybe it goes well and there's some sort of slowdown agreement, but either way, it's gonna get decided in the next couple of years.
Based on my conversations with people at Anthropic, this is a pretty well-understood thing inside the company. They think that they need to build the most powerful AI to make sure that this transition goes well. Do you think that there's any inherent problems with that view?
I think there’s some pretty obvious problems with that view. But I would like to first say that I'm very sympathetic to it, in that I think Anthropic is far and away the most responsible player in the space. Having worked at both OpenAI and Anthropic, there is a night-and-day difference in the extent to which they're taking the situation seriously.
Can you say how?
To give an example, executives at OpenAI won't give you their exact pictures for what the world will look like. They'll never say, like, “This is exactly why we're doing this, and this is the way the world will look.” They won’t give concrete predictions. Anthropic will have executives and leadership making very clear predictions and discussing, like, details of company strategy with the whole company.
And part of the reason they can do this is because everyone at Anthropic is treating this like we're on war footing. They don't leak anything—nothing ever leaks from Anthropic [Editor’s note: Some stuff has leaked]. OpenAI has leaks every day. Anthropic has no leaks because, again, the people inside it are treating this like a serious mini Manhattan Project. Except obviously the difference is that [Anthropic] doesn’t have a government mandate. This is a private company acting like it's the Manhattan Project.
Do you think that the world should trust Anthropic to run this mini Manhattan Project?
I think no private company should. I think Anthropic is doing their best, and they're doing a very good job, but the structural necessity of the race is that, in the future, they will have to cut corners, so they will have to make trade-offs between rigor and safety, because they're racing against competitors like OpenAI and China. They're begging for regulation. Many [Anthropic leaders] have gone on the record saying they want to be regulated, and the reason they want to be regulated is because they are scared of the race that they're in.
It’s not really about whether we should trust [Anthropic]. We need to step in and ensure that the race isn't happening because [Anthropic] can't really trust themselves in the context of the race.
Do you think Anthropic is already cutting corners?
No, not yet. It's not cutting any corners. What I'm saying is that when things speed up in the next year, they'll have to if they want to to stay competitive.
So, people are kind of skeptical of this claim that AI will kill us all. I think the main question I get from people is, well, how will AI actually kill us all? And is that the right question to be asking?
I think it's a really natural question to ask, because, as I've said many times, this stuff sounds like science fiction. But the classic example is, like, synthesizing a new virus or taking down critical infrastructure by doing some sort of hacking spree—the latter might be less likely to kill literally everyone.
But I think the more important question than that is, like, who are the people saying this is possible? And if you look at the record, you have many leading fathers of machine learning and artificial intelligence saying extinction is a possibility. You have both Dario Amodei and Sam Altman in the last five years.
One thing that could be pretty interesting is if people were to get these executives on the record again and ask them to give an actual probability for extinction in the next decade and see what number they come up with. Because this view is shared among many of the people that are thinking about this seriously, but they don't often express it so bluntly.
So you’ve clearly garnered a lot of attention. There's a lot of talk in the industry about a pause or a mechanism to pace AI development, or some sort of regulation on frontier AI models from the government. What do you think needs to be done now?
A baby step would be some sort of agreement between OpenAI and Anthropic as the leading Western labs. [They would need] some sort of neutral understanding that they won't immediately go into recursive self-improvement in the next year. These are the two places where it’s likely to happen right now.
The problem with the pause is you've got China. So, what you really ultimately want is some sort of international agreement—international pacing, you might call it—which will require international coordination and understanding where all the compute in the world is. Maybe even some sort of international CERN-like institution. There are a lot of great proposals on this. AI 2040, if you’ve heard of it, is like the successor to AI 2027.
All of these ideas involve quite painful government intervention. But when you internalize the risk, and when you actually get executives on the record and they admit how dangerous this stuff is, I don't see any other option than treating the computers that run AI as a dangerous resource. We need to know who has which computers in the same way that we need to know who has which nuclear materials.
Jacob, is there a reason that kept you showing up to work at OpenAI and Anthropic every day for the last few years? Do you feel like there is a path forward here where everything is fine?
I really want to see the benefits of this technology. For example, it's also not hyperbole that we could cure cancer. Maybe you saw the achievement yesterday on Navier-Stokes, the novel mathematical problem. There really is no reason that we can't transfer that to scientific domains like biology and make breakthroughs on questions we've struggled with for a long time.
It feels like we're sitting on the doorstep of ridiculous abundance if we can make this technology go right. The main issue, right now, is entering that world of abundance with moderation. We can just get so excited, rush straight into the self-improving thing in the next couple of years, and blow it all up.
If we're careful and sensible, there’s a lot of ways to get tremendous value from these models.
What's the feedback been like from other AI researchers about this?
I've not been talking to that many, but I think a lot of my colleagues at Anthropic are genuinely happy that this has gained so much traction. Again, a lot of them are just quite pessimistic about the world waking up. Like, they're just worried that no one is really going to come and save them from the race.
I think it’s quite surprising that many of them are genuinely happy that a tweet that criticizes them is getting lots of traction. It should make people think more about the type of people that are at Anthropic. They're a very well-meaning, interesting, and kind of weird set of people.
I'm sure the next couple days will be crazy, but after things settle down, have you figured out what's next for you?
I mentioned AI 2027 and AI 2040 before. I think these are really interesting predictions of the way the future looks. They're not done by any big labs; these were done externally by researchers that just cared about predicting what the world looks like. These have been very good resources for me and a lot of other researchers for trying to think through the implications of what you're doing.
So I expect to do something like this—some sort of independent commentary on where things are going, at least for the short term before we get some sort of pacing agreement or slowdown. [If] it looks like we're not going to rush straight into recursive self-improvement, maybe I’d consider going to an auditing agency or some third-party regulation or transparency agency. But for now, I think I'm just gonna try and take stock of where the situation is.
Update 9/9/26 7:40pm ET: This story has been updated with a statement from Anthropic.
Comments
Back to top

连线杂志AI最前沿

文章目录


    扫描二维码,在手机上阅读