OpenAI 推出 GPT-Red 以测试 AI 模型安全性

qimuai 发布于 阅读:52 一手编译

OpenAI 推出 GPT-Red 以测试 AI 模型安全性

内容来源:https://aibusiness.com/generative-ai/openai-ugpt-red-tests-ai-model-safety

内容总结:

谷歌云赞助报道:如何选择首个生成式AI应用场景

企业布局生成式AI,应优先聚焦于能够改善人类信息交互体验的领域。与此同时,AI模型的安全性正成为关键议题。

AI红队测试:从人工到人机协同

传统红队测试(人工模拟攻击以发现漏洞)已不新鲜,但结合人类与AI共同测试新模型安全性却是前沿做法。企业仍需确保所选模型符合自身业务与安全流程。

OpenAI发布GPT-Red:自动化红队模型

本周三,OpenAI正式推出GPT-Red——一款自动化红队测试模型。该模型用于内部模拟“劫持”或“越狱”OpenAI模型,尤其擅长提示注入攻击(通过恶意指令操控AI)。OpenAI利用GPT-Red生成的攻击案例来训练生产模型,并用于评估已有AI软件的安全性。

应对前沿AI安全焦虑

GPT-Red的推出,直接回应了业界对Anthropic Claude Fable、Mythos等强大模型可能暴露系统漏洞的担忧。甚至特朗普政府已出台自愿隔离政策,允许联邦政府在模型公开发布前先行评估。

独立AI分析机构Tekonyx创始人Sid Nag指出:“前沿AI系统正变得日益关键,需要更标准化、独立的评估流程。”他强调,部署前进行红队测试,能帮助企业做出最佳模型选择决策,“虽然无法完全消除滥用或突现行为的担忧,但透明且可重复的测试能建立信心。”

行业共识:人机协同成标配

Forrester分析师Mike Gualtieri表示,OpenAI并非唯一采用人机协同红队测试的厂商。Anthropic同样高调推行其“前沿红队”,利用AI模拟攻击后再公开发布模型。他认为这应成为模型部署前的常规实践。

企业需警惕:安全测试非万能药

但Gualtieri也提醒,红队测试存在副作用:“一人眼中的漏洞,可能是另一人眼中的隐私或自由。”安全措施本身也可能存在脆弱点。因此,企业不应依赖厂商的自动化红队测试而放弃自身的尽职调查。

“如果焦点始终放在网络安全上,那或许是好事,但企业仍需保持警惕。”Gualtieri补充道,供应商的安全措施不应干扰企业按自身需求使用模型的方式。

中文翻译:

由谷歌云赞助
选择你的首个生成式AI应用场景
要开始使用生成式AI,首先应聚焦于能够改善人类与信息交互体验的领域。

尽管红队测试已是常规操作,但利用人类与AI共同测试新模型的安全性仍属创新之举。企业仍需确保所选模型与其业务及安全工作流程相契合。

OpenAI于周三发布了自动化红队测试模型GPT-Red。这家初创公司的最新成果表明,AI模型正成为需要重点防护的重要工具——既要防止其落入不当用途,也要保障企业数据和工作流程的安全。

据OpenAI称,GPT-Red是其测试GPT-5.6 Sol等最新模型安全性的集大成之作。该模型内部用于尝试劫持或破解OpenAI模型,专精于提示注入攻击——即黑客通过在提示词中植入恶意指令来操控AI模型或代理。OpenAI表示,他们利用GPT-Red生成的攻击训练生产模型,并针对实际运行中的AI软件进行安全评估。

GPT-Red的推出源于近期对Anthropic公司Claude Fable和Mythos等模型的担忧——这些模型过于强大,可能暴露漏洞或降低系统入侵门槛。这种忧虑甚至促使特朗普政府引入自愿隔离政策,以便联邦政府在新模型公开发布前有充分时间进行评估。

"GPT-Red承认前沿AI系统正变得愈发关键,"独立AI分析机构Tekonyx创始人兼CEO、技术专家Sid Nag表示,"这些系统需要更标准化、独立化的评估流程才能实现广泛部署。"

Nag补充道,在部署前对AI模型进行红队测试、威胁排查及破解漏洞评估,能使企业针对模型选型做出最优运营决策。

"这不会完全消除对滥用或突发行为的担忧,"Nag说,"但透明化与可复现测试能建立信任。"

Forrester分析师Mike Gualtieri指出,OpenAI并非唯一采用人机协同方式进行红队测试的厂商。"Anthropic同样高度重视引导模型行为符合预期并确保安全性。"他补充道,这似乎应成为模型部署前的标准流程。

Anthropic此前也公开披露其前沿红队团队,以及在新模型公开发布前使用AI模拟攻击的举措。

不过Gualtieri认为,红队测试对企业客户可能存在弊端。

"并非全是优点——某人眼中的漏洞,在另一个人看来可能是隐私或自由权利。"他表示,某些安全措施本身也可能存在脆弱性。因此,企业在评估最适合自身工作流程的模型时,不应将厂商的自动化红队测试视为放弃自主尽职调查的理由。

"若能聚焦网络安全,这或许是好现象,但企业仍需保持警惕。"Gualtieri补充道,厂商采取的安全措施也不应干扰企业对模型的实际使用方式。

英文来源:

Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
While red teaming is standard practice, using humans and AI to test the security of new models is novel. Enterprises should still ensure the model they use aligns with their business and security workflows.
OpenAI on Wednesday launched GPT-Red, an automated red teaming model. The startup's latest release reflects how AI models are becoming important tools that need to be safeguarded to prevent them from falling into the wrong hands and to keep enterprise data and workflows protected.
GPT-Red is the culmination of OpenAI’s efforts to test the safety of its recent models, such as GPT-5.6 Sol, according to the company. The model is used internally to try to hijack or jailbreak OpenAI models; it specializes in prompt injection, a method in which a hacker manipulates an AI model or agent by inputting malicious instructions into a prompt. OpenAI said it used the attacks that GPT-Red generated to train its production models and evaluate them against AI software in production.
GPT-Red is a response to the recent fear about models such as Anthropic’s Claude Fable and Mythos being so powerful they could possibly expose vulnerabilities or make it easier to hack into systems. The concern is so great, the Trump Administration introduced a voluntary quarantine policy to allow the federal government time to evaluate models before they’re released to the public.
"GPT-Red acknowledges that frontier AI systems are becoming more critical," said Sid Nag, tech expert and founder and CEO of independent AI analyst firm Tekonyx. "They require more standardized and independent evaluation to go and deploy these things broadly."
Nag added that red teaming and testing AI models for potential threats and trying to jailbreak them to test for vulnerabilities before deployment means enterprises can make the best operational decisions about which models to use.
"It will not eliminate all concerns around misuse or emergent behaviors," Nag said. "But transparency and repeatable testing build confidence."
OpenAI is not the only vendor combining humans and AI to red team new models, said Mike Gualtieri, an analyst at Forrester. "Anthropic does make a big deal of trying to make the model behave in a way that they want it to behave and be safe," Gualtieri said. He added that this seems to be a customary practice that should be followed before a model is deployed.
Anthropic has also been open about its Frontier Red Team and its use of AI to simulate attacks on new models before releasing them to the public.
Nevertheless, there could be a downside to red teaming for enterprise customers, Gualtieri added.
"It's not all good because what one person thinks is a vulnerability, another might think is privacy or think is their freedom," he said. He added that some safety measures could be vulnerable as well. Therefore, enterprises should not see a vendor's use of automated red teaming as a reason to forgo their own due diligence when evaluating models that best fit their workflows.
"If the focus remains on cybersecurity, that’s probably a good thing, but enterprises should still be wary of that," Gualtieri said. He added that the safety measures vendors take also shouldn't interfere with how enterprises want to use the model.

商业视角看AI

文章目录


    扫描二维码,在手机上阅读