人工智能护栏如何阻碍进攻性网络安全研究人员的工作

qimuai 发布于 阅读:49 一手编译

人工智能护栏如何阻碍进攻性网络安全研究人员的工作

内容来源:https://techcrunch.com/2026/07/23/how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers/

内容总结:

AI巨头设防“护栏”反成枷锁,网络安全攻防两难引争议

过去数月,人工智能巨头们精心设计了审查计划与严格的安全护栏,旨在限制恶意黑客对其模型的使用。然而,这些限制如今却正阻碍合法网络防御者以及进攻性网络安全研究员的正常工作。

今年6月,美国政府以“可能被用于发起恶意网络攻击”为由,对Anthropic公司备受瞩目的AI模型“Mythos”与“Fable”施加了出口管制。尽管具体原因存在争议,但事实是,Anthropic曾反复将Mythos宣传为某种“末日网络武器”,仅允许经过严格审查的用户使用,并配有严密护栏。(目前相关限制已部分解除。)

这种“设卡”做法并非个例。OpenAI与Anthropic均推出了针对网络安全研究员的“可信访问”或“验证计划”,申请通过后可使用限制更少的模型。但该做法遭到广泛批评。知名安全研究员马克·多德直言,由大公司“随意决定什么安全、什么不安全”,令人感到不适。

核心问题在于,安全护栏在阻断恶意攻击的同时,也伤及防御者。安全咨询集团NCC Group首席科学家克里斯·安利指出,要求AI模型尝试利用漏洞,是确认漏洞真实性的关键步骤。但若护栏直接拒绝回答,等于帮了攻击者,害了防御者。他强调,修复代码与发现漏洞本质上是同一工具的两面,无法被简单剥离。当AI模型拒绝配合时,研究员们只能倒向毫无护栏的开源模型。

一些进攻型安全公司明确表示,他们会刻意避免将敏感漏洞发现工作上传至云端模型,以防数据泄露或被用于后续训练,而是选择本地运行的开源模型。CrowdFense首席技术官保罗·斯塔格诺批评AI公司“本质上像对待需要保姆的小孩一样对待客户”。

更有研究员反映,其所在公司因未加入Anthropic的验证计划,导致工具几乎毫无用处。安全公司RemoteThreat首席执行官克里斯·汤普森称,即便在审查计划内,护栏也极不稳定,“你大部分时间在跟模型讨价还价,而不是专注于核心安全工作”。他警告,这正迫使负责任的美国研究员转向GLM等中国开源模型,从而“被推向外国系统”。汤普森呼吁AI实验室开放程序、提供负责任的访问,并追究滥用者责任,否则防御方将在AI竞赛中落败。“一场前所未有的攻击浪潮即将来临,但那些真正想有所作为的安全公司正被死死扼住喉咙。”

中文翻译:

数月以来,人工智能巨头们设计了经过特别审查的准入程序与严格护栏机制,以限制恶意黑客使用其模型。但这些限制如今不仅阻碍了合法网络防御者的工作,也影响了进攻性网络安全研究人员的行动。
今年六月,美国政府以出口管制手段限制了Anthropic备受瞩目的AI模型Mythos与Fable。此举至少部分源于一份报告,该报告声称能够绕过模型为防止用户构建并实施恶意网络攻击而设置的护栏机制。
无论此事件是否真正源于对越狱攻击的担忧,事实是Anthropic一再将Mythos宣传为某种"末日网络武器",声称其只能交付经过严格审查的用户——即便对这类用户也需配备严密的护栏。(针对Fable 5和Mythos 5的出口管制现已解除。Fable 5于7月1日恢复公共访问;而Mythos 5作为政府审查流程的一部分,仅重新向通过审查的美国机构开放。)
这种设限做法并非Mythos独有。Anthropic的其他模型以及OpenAI均向网络安全研究人员提供申请审核通道——获批后即可获取网络安全限制更少的模型:例如OpenAI的"可信网络安全访问"计划与Anthropic的"网络安全验证计划"。
这些护栏机制饱受批评,尤其遭到那些致力于在系统中发现未知漏洞、并在黑客攻击前设计利用方法的研究人员的反对。
知名安全研究员马克·多德近期在网络安全播客中表示:"这些随机的大型企业竟能对安全领域什么是安全的、什么是不安全的做出武断决定,这让我感到不适。"多德数十年来致力于发现并向西方政府出售"零日漏洞"——即此前未知的软件缺陷及其利用程序——而非向软件制造商报告以修复漏洞。政府之所以为这类漏洞支付溢价,正是因为它们始终处于未修复状态,这对情报行动极具价值。
多德承认其工作性质可能带来偏见,但持有类似观点者不在少数。多位从事进攻性网络安全工作(即主动探测系统弱点)的人士向TechCrunch描述了其使用AI工具及应对护栏机制的方式。
安全咨询巨头NCC集团首席科学家克里斯·安利表示,要求AI模型尝试利用某个漏洞是确认其是否值得修复的关键步骤。但若护栏机制导致模型直接拒绝回答,反而会损害防御者的利益。
安利指出:"这正是进攻与防御及护栏机制的交汇点。因为'修复这段代码'这类指令既是防御的核心机制,也是发现代码库关键漏洞的路线图。同一工具同时具备攻防属性,两者根本无法剥离。"他比喻道:"就像锤子,没有锤子无法建造房屋。它毫无疑问是工具,但也不可避免地是武器。"
当他和同事遭遇此类阻碍时,有时会转向完全无护栏的开源AI模型。
著名漏洞发现与交易公司CrowdFense的首席技术官保罗·斯塔尼奥赞同多德的观点,认为AI公司通过审查程序和护栏"本质上将客户视为需要保姆的孩童"。斯塔尼奥表示,他和同事仅将前沿模型用于逆向工程,而避免使用AI辅助发现漏洞或构建利用程序——因为将此类工作输入云端模型可能导致敏感漏洞数据泄露,或被纳入未来训练数据。为此,他们选择在本地运行开源模型,从而避免依赖模型之外的数据共享。
专注于零日漏洞发现与利用程序开发的安全研究员朱塞佩·卡利表示,护栏并未妨碍其工作。他并未将AI用于进攻性操作,而是用于初步逆向工程、理解分析目标代码及构建辅助工具。他认为AI工具能加速流程,使其专注于漏洞发现:"我仍希望亲自掌控漏洞发现与武器化过程,即使明天所有护栏都被取消,这也不会改变。我对自己的漏洞充满保护欲,且对这场游戏的热爱程度足以让我拒绝让模型代劳。"
某智能手机组件制造商研究员(因未获授权与媒体沟通而要求匿名)透露,其公司未加入Anthropic的CVP计划,导致其工具因护栏过于严格而难以用于漏洞发现。该人士称:"一旦模型察觉到我们从事安全相关工作,就会直接停止响应,完全不可用。"
网络安全公司RemoteThreat首席执行官、进攻性安全与AI主题会议"攻防AI大会"创始人克里斯·汤普森指出,根据其使用前沿AI模型的经验,护栏机制每天的表现不一致且效果各异——即便在Anthropic和OpenAI审核计划提供的宽松框架内也是如此。
他认为:"实际影响在于,用户花费大量时间与模型'讨价还价',而非专注于核心安全任务。你本应分析漏洞并推演其可利用性,却不得不耗费精力探究结果不一致的原因,或是模型为何过度净化输出内容。"
汤普森表示,研究人员因此被迫转向GLM等中国开源模型——这些模型可免费下载并在本地运行,无需审查或使用限制。他说:"负责任的研发者正被美国监管系统推向外国的系统。我认为设置这些护栏弊大于利。"
汤普森呼吁前沿AI实验室放宽限制,开放准入计划并提供负责任的使用权限,同时追责滥用工具者。否则,防御者将在AI竞赛中落败:"一场巨大的风暴正在逼近,我们将遭遇前所未有的速度与规模兼有的攻击浪潮。但那些试图有所作为的合规安全咨询公司与研究人员,此刻正遭受压制。"

英文来源:

For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.
In June, the U.S. government slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.
Regardless of whether the incident was really motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government’s review process.)
That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and — if approved — access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program.
These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”
Dowd has spent decades finding and selling “zero days” — previously unknown software flaws and the exploits that take advantage of them — to Western governments, rather than report them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.
Dowd admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity — they proactively probe systems for weaknesses — described to TechCrunch how they use AI tools and deal with their guardrails.
Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.
“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”
When he and his colleagues run into such a roadblock, they sometimes fall back on open-source AI models that come with no guardrails at all.
Paolo Stagno, the chief technology officer at CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies “essentially treat customers like children who need babysitting” with their vetted programs and guardrails.
Stagno said he and his colleagues do use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said.
Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event — said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That’s true even inside the looser boundaries of Anthropic and OpenAI’s vetted programs.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
As a consequence, researchers rely on or get pushed toward Chinese open-source models like GLM — freely downloadable models that can be run locally with no vetting or usage restrictions — said Thompson.
“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”
Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and also hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”

TechCrunchAI大撞车

文章目录


    扫描二维码,在手机上阅读