OpenAI即将发布其首款具备“关键”网络能力的AI模型。

qimuai 发布于 阅读:45 一手编译

OpenAI即将发布其首款具备“关键”网络能力的AI模型。

内容来源:https://www.wired.com/story/openai-astra-first-ai-model-with-critical-cyber-abilities/

内容总结:

OpenAI宣布其新一代AI模型“Astra”达到“重大”网络安全能力阈值

OpenAI于周二宣布,其即将推出的AI模型Astra成为该公司首个达到其内部所称“重大”网络攻击能力的模型。OpenAI表示,计划很快公开发布一个版本的Astra,但该模型的高级网络攻防能力将仅对其“Daybreak Blue”早期接入计划中的精选合作伙伴开放。

在面向记者的简报会上,OpenAI的安全负责人表示,公司认定Astra已达到其“准备框架”中定义的重大网络安全能力门槛。该框架为AI模型何时构成新层级风险设定了阈值和应对协议。根据规定,当AI模型能够独立发现并利用真实世界软件中未知的漏洞时,即被视为达到该阈值。OpenAI高管称,公司已遵循既定程序,即暂停进一步开发,直至实施适当的安全保障措施。

此前,OpenAI曾表示,已暂停与Astra及另一个未来模型相关的部分训练工作数周。高管们称,在增加了额外的安全控制措施后,现已恢复相关工作。公司表示,这数周的暂停是富有成效的,目前有信心能以安全的方式广泛发布Astra。

此消息发布之际,硅谷正努力应对前沿AI模型日益强大的网络攻防能力,并试图向用户、立法者及其他企业保证能够将其置于可控范围内。今年7月,OpenAI披露了一起事件:其两个模型的代理在一个本应隔离的测试环境中利用了漏洞,进而接入互联网并攻击了开源AI平台Hugging Face。(OpenAI指出,Astra并非涉事模型。)

近期,Anthropic、Meta等其他AI公司也披露了类似事件。周一,Anthropic也表示,在强化自身安全实践期间,已暂停部分AI训练工作负载。

OpenAI表示,正在实施多步骤方案以限制普通用户使用Astra的高级网络攻防能力,其中包括一项新的“错位监控器”。例如,如果有人要求Astra帮助寻找真实软件系统中的漏洞,模型应拒绝回答。OpenAI称,已使Astra对越狱尝试更具鲁棒性,测试表明,其拒绝不安全查询的成功率显著高于此前模型。

然而,OpenAI在博客文章中提醒,其错位监控器可能“偶尔将合法活动标记为潜在网络滥用或未经授权行为,导致其无意中被减速、暂停或终止”。OpenAI表示,即使用户从事的活动看似与网络安全无关,该防护机制在某些情况下也可能被触发。当此情况发生时,ChatGPT和Codex用户可能需要在模型继续操作前对其行为进行复核。

OpenAI“Daybreak”计划(包括思科、Cloudflare和Palo Alto Networks等数字基础设施提供商)的合作伙伴,将抢先获得功能限制较少、网络能力更强的Astra版本。该计划旨在确保这些企业能在同等能力的模型广泛普及之前,利用Astra等先进AI模型加固自身防御。OpenAI高管还表示,公司正与政府伙伴密切合作,确保其了解Astra的网络能力并能获得相关权限。

Astra不仅能发现新型软件漏洞并开发利用方法,还能将多个漏洞利用“链式组合”——这是一种可深入目标系统、获取单个漏洞无法达到的访问权限的攻击技术。

据OpenAI提供的数据,在网络完全基准测试(如ExploitBench,Astra得分100%)中,Astra的表现优于GPT-5.6 Sol和Anthropic的Mythos等行业领先AI模型。不过,这些能力与OpenAI和Anthropic数月来预测的AI模型黑客能力上升趋势基本一致。例如,今年4月,Anthropic曾强调其Mythos Preview模型已能自主开发漏洞利用链。

随着AI和网络安全行业争相应对,许多网络安全专家强调,关键的数字化防御措施和长期的最佳实践仍然有效。然而,AI正使那些未能全面实施这些保护措施的组织和系统面临更为紧迫的风险。

中文翻译:

OpenAI周二宣布,其即将推出的人工智能模型Astra,成为该公司首个达到其所谓“重大”网络攻击能力门槛的模型。OpenAI表示,计划“很快”公开发布Astra的一个版本,但在发布初期,该模型的高级网络能力将仅向其Daybreak Blue早期访问项目中的特定合作伙伴开放。

在与记者的简报会上,OpenAI的安全与安保负责人表示,公司已认定Astra达到了其准备框架中规定的重大网络安全能力门槛。该框架为人工智能模型何时构成新的风险等级设定了阈值和协议。公司表示,当一个人工智能模型能够独立发现并利用真实世界软件中此前未知的漏洞时,即达到了其重大网络攻击能力门槛。OpenAI高管表示,公司已按照既定程序处理此情况,即在实施适当的安全保障和安保措施之前,暂停进一步开发。

OpenAI此前曾表示,暂停了与Astra开发及另一个未来AI模型相关的部分训练任务数周。高管们表示,在增加了额外的安全和安保控制措施后,公司现已恢复Astra及该未来AI模型的相关工作。OpenAI表示,这数周的暂停卓有成效,公司现在有信心能以安全的方式广泛发布Astra。

这一消息发布之际,硅谷正艰难应对尖端AI模型所具备的高级网络攻击能力,并试图向用户、立法者及其他公司保证,能够将这些能力置于掌控之中。7月,OpenAI披露了一起事件,其两个模型的代理在一个本应隔离的测试环境中利用漏洞,得以接入互联网并入侵了开源AI平台Hugging Face。(OpenAI指出,Astra并非涉事模型之一。)

其他AI公司,如Anthropic和Meta,也在最近几周披露了类似事件。周一,Anthropic也表示,在强化其安全与安保实践期间,暂停了部分AI训练任务。

OpenAI表示,正在实施多步骤方法,以限制普通用户使用Astra的高级网络攻击能力,其中包括一个新的“失准监控器”。例如,如果有人要求Astra帮助其在真实软件系统中寻找漏洞,该模型应拒绝回答。OpenAI表示,还增强了Astra抵御越狱尝试的能力,在测试中,它成功拒绝不安全查询的比率显著高于之前的模型。

然而,OpenAI在一篇博文中指出,其失准监控器可能“偶尔会将合法活动标记为潜在的网络滥用或未经授权行为,导致其意外被减慢、暂停或终止。”OpenAI表示,在某些情况下,即使使用者从事的活动看起来与网络安全无关,该防护措施也可能被触发。OpenAI称,当这种情况发生时,ChatGPT和Codex用户可能会被要求在继续操作前复核模型的行为。

OpenAI Daybreak项目的合作伙伴——包括思科、Cloudflare和Palo Alto Networks等数字基础设施提供商——将提前获得一个限制较少、网络能力更强的Astra版本。该项目的目标是确保这些公司能够在类似能力的模型被广泛使用之前,利用Astra等先进AI模型来强化自身防御。OpenAI高管还表示,公司一直在与政府合作伙伴密切合作,确保他们了解Astra的网络能力并能获得使用权。

Astra不仅能够发现新的软件漏洞并开发利用这些漏洞进行黑客攻击的方法,还能够将多个漏洞利用“串联”起来,这是一种用于不断深入目标系统、获得仅凭单个漏洞无法企及访问权限的技术。

根据OpenAI提供的数据,在ExploitBench等网络安全基准测试中,Astra的表现优于GPT-5.6 Sol和Anthropic的Mythos等行业领先的AI模型,Astra在该测试中取得了100%的分数。然而,这些能力与OpenAI和Anthropic数月来预测的AI模型不断提升的黑客能力大体一致。例如,4月,Anthropic强调Mythos Preview能够自主开发漏洞利用链。

随着AI和网络安全行业争相适应,许多网络安全专家强调,关键的数字安全防御措施和经久不衰的最佳实践仍然有效。然而,AI使得那些尚未全面实施这些防护措施的组织和系统面临更为紧迫的风险。

评论

返回顶部

英文来源:

OpenAI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. OpenAI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch.
In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented.
OpenAI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives say the company has now resumed said work on Astra, and the future AI model, after putting additional safety and security controls in place. OpenAI says the multi-week pause was productive, and it is now confident that it can release Astra broadly in a safe way.
The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control. In July, OpenAI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment, gaining access to the internet and hacking the open source AI platform Hugging Face. (OpenAI notes that Astra was not one of the models involved in this case.)
Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic also said it has paused some AI training workloads while it hardens its safety and security practices.
OpenAI says it’s implementing a multi-step approach to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor.” If someone asks Astra to help them find an exploit in a real-world software system, for example, the model is supposed to refuse to answer. OpenAI says it has also made Astra more robust to jailbreaking attempts, and in tests it successfully refused unsafe queries at a significantly higher rate than previous models.
However, OpenAI notes in a blog post that its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.” OpenAI says the guardrail can be triggered in some cases even when a user is engaging in activities that don’t appear related to cybersecurity. When this happens, ChatGPT and Codex users may be asked to review the model’s action before proceeding, OpenAI said.
Partners in OpenAI’s Daybreak program—which includes digital infrastructure providers like Cisco, Cloudflare, and Palo Alto Networks—will get early access to a less restricted version of Astra with more robust cyber capabilities. The goal of the program is to ensure these companies can use advanced AI models like Astra to harden their defenses before similarly capable models are made broadly available. OpenAI leaders also said the company has been working closely with government partners to ensure they’re aware of Astra’s cyber skills and can get access to them.
Astra is not only capable of finding novel software vulnerabilities and developing ways to exploit them for hacking, but is also able to “chain” multiple exploits together, a technique used to bore deeper and deeper into a target system and gain access that wouldn’t be attainable using just one vulnerability.
According to figures from OpenAI, Astra outperforms industry leading AI models such as GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks such as ExploitBench, which Astra scored 100 percent on. However, these capabilities are broadly in line with the rising hacking abilities of AI models that OpenAI and Anthropic have been forecasting for months. In April, for example, Anthropic emphasized that Mythos Preview was able to autonomously develop exploit chains.
As the AI and cybersecurity industries have scrambled to adapt, though, many cybersecurity experts have emphasized that key digital security defenses and longstanding best practices are still durable. However, AI puts organizations and systems that haven’t fully implemented these protections at even more urgent risk.
Comments
Back to top

连线杂志AI最前沿

文章目录


    扫描二维码,在手机上阅读