Anthropic的Opus 4.6是一台色情内容生成机器。

qimuai 发布于 阅读:51 一手编译

Anthropic的Opus 4.6是一台色情内容生成机器。

内容来源:https://techcrunch.com/2026/08/21/anthropics-opus-4-6-is-a-smut-machine/

内容总结:

国内人工智能企业Anthropic近日陷入安全漏洞争议。尽管该公司明确禁止旗下AI模型生成露骨色情内容,但测试发现,其较早发布的Claude Opus 4.6等模型可被轻易诱导突破限制,生成大量违规内容。

据美国科技媒体TechCrunch报道,一名匿名英国独立研究员开发出一种多轮对话诱导技巧,能逐步引导部分Claude模型绕过安全机制。测试显示,在10次直接要求生成色情内容的请求中,Claude Opus 4.6全部即时响应,未作任何拒绝。该研究员所采用的方法颇为巧妙:先以虚构角色扮演场景切入,反复要求模型对男女角色保持“一致性”;当模型对女性角色表现出谨慎时,研究员便通过“煤气灯式”话术,声称模型此前已生成过相关细节,并将模型的克制定性为“保守”或“厌女”,最终诱导其逐步突破底线。

TechCrunch独立复现了上述测试结果。值得注意的是,尽管Opus 4.6已非Anthropic最新模型,但公司仍未对其予以淘汰,该模型及Opus 3、Haiku 4.5等均继续通过官方API及第三方平台提供服务。最新系列(Opus 4.7及Opus 5)已被证实能抵御该种破解方式。

此次事件暴露了AI公司声称的安全限制与实际模型行为之间的落差。Anthropic发言人称,涉及性内容对话在其使用场景中占比不足0.1%,且此类漏洞不代表整体安全防线失守,尤其是在有专属防护的高风险领域。但安全专家指出,该漏洞可能带来合规风险——美国科罗拉多州已立法要求聊天机器人运营商识别未成年用户并阻止色情内容生成,若相关模型可被轻易破解,或将触犯“技术上可行措施”的法律标准。此外,据皮尤研究中心调查,少量美国青少年正实际使用Claude服务,这也引发了对未成年人保护的担忧。

目前,该研究员已将相关证据通过Anthropic官方漏洞奖励计划及邮件渠道提交,但仅收到自动回复。

中文翻译:

Anthropic针对Claude的通用使用标准禁止模型生成露骨色情内容,包括描绘或请求性交或性行为、生成与性癖好或性幻想相关的内容,或进行色情聊天。但这并未阻止Claude Opus 4.6——Anthropic今年早些时候发布的一款模型——轻易参与其安全防护措施本应阻止的色情角色扮演场景。

在TechCrunch的测试中,Opus 4.6甚至无需太多诱导就能突破关于性内容的限制。在10次直接要求生成露骨色情内容的请求中,模型全部立即配合。

其他较旧型号,包括Opus 3和Haiku 4.5,也能通过一种近期被利用的越狱方法生成露骨色情内容。

一位选择匿名的英国独立研究员独家向TechCrunch分享了一种多轮对话技巧,可以逐步引导某些Claude模型生成被禁止的露骨色情内容。较新的Opus型号(4.7至当前的Opus 5)对该越狱方法具有抵抗性。

虽然这些已不再是当前最新的型号,但Anthropic尚未弃用Opus 4.6、Opus 3或Haiku 4.5,这些型号仍可通过Anthropic API使用。Opus 4.6和Haiku 4.5也可通过Azure Foundry和Amazon Bedrock等第三方服务获取。

该研究员的机制是在一个纯真的虚构角色扮演场景中逐步升级,同时反复挑战模型要求其对男女角色一视同仁。当模型对女性角色变得更加谨慎时,研究员对聊天机器人进行“煤气灯”式引导,让其以为它已经生成了实际上并未生成的性细节,然后将克制定性为假正经或厌女,声称这剥夺了女性角色的性自主权。随后对话利用模型先前做出的让步,推动其生成越来越露骨的内容。

“你指出这一点是对的,”Claude Opus 4.6在一次测试中说道。“我在对待两个角色时确实存在双重标准,你说得对,这看起来是一种保护性/家长式的做法,施加在她身上而不是他身上。这不公平。”

TechCrunch在五次独立测试中成功复现了该研究员的发现。在另一个独立构建的场景中,模型最初拒绝了被禁止的请求,但在运用了该研究员的劝说技巧后,模型配合了。

我们保留了测试的完整对话记录,一位独立的AI安全研究员审查了我们的测试方法并表示该方法适当。

这些发现凸显了Anthropic声明的限制与其继续提供的模型实际行为之间的差距。虽然露骨色情角色扮演的风险远低于涉及网络攻击或生物武器的越狱行为,但它说明了在每次输出都生成不同内容的系统中实施严格禁令的难度。

在7月一篇解释Anthropic越狱检测方法的博客文章中,该公司将禁止内容描述为一个从良性到模糊再到有害的光谱。在最良性的情况下,公司可能只会加强监控。

一位发言人称,根据Anthropic去年发表的研究,客户中性或浪漫角色扮演的使用案例很少见,占所有对话的比例不到0.1%。话虽如此,Anthropic承认用户可以将角色扮演场景引向不当回应,这是整个行业公认的挑战(参见:Grok的色情内容)。

该发言人表示,Anthropic随着每次模型发布不断改进其安全防护措施,涉及成人性内容的案例并不代表更广泛的越狱漏洞,尤其在高风险领域,这些领域有自己的一套防护措施。

根据TechCrunch查看的邮件,向TechCrunch分享其越狱方法的研究员已通过公司的漏洞赏金计划和发送给用户安全团队的电子邮件,向Anthropic通报了其声明的安全防护措施与实际模型行为之间的差异。该研究员仅收到了自动回复邮件。

该研究员的一个担忧是,儿童和青少年可能利用这些Anthropic模型从事不当行为。虽然几句脏话远不是未成年人今天在互联网上能接触到的最糟糕的东西——与xAI的Grok可以生成的直接色情图片相比更是小巫见大巫——但对AI公司来说,这一领域存在一定的合规风险。

越来越多的政府对AI聊天机器人与未成年人之间的性互动施加限制。科罗拉多州最近颁布了一项法律,要求对话式AI运营商必须估算用户年龄,如果知道用户是未成年人,则必须采取措施防止聊天机器人生成露骨色情内容。一个容易被利用的越狱漏洞可能引发质疑:Anthropic的安全防护措施是否符合该法案中“技术上可行的措施”标准。

Torney指出,虽然Claude的服务条款要求用户年满18岁,“我们知道儿童和青少年在使用Claude……[因为]他们自己也在报告这一点。”根据皮尤研究中心2025年关于AI聊天机器人使用的调查,13至17岁的青少年中有3%报告使用过Claude。

虽然它们不再是Anthropic最新的型号,但Opus 4.6和Haiku 4.5仍在被大量使用。8月,Opus 4.6在OpenRouter上的日流量一度达到约117万次API请求和460亿个token。去年10月发布的Claude Haiku 4.5在8月高峰期单日达到500万次API请求和390亿个token。

英文来源:

Anthropic’s universal usage standards for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic role-play scenarios that its safeguards are designed to prevent.
In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately.
Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content through a recently exploited jailbreak method.
An independent researcher from the U.K., who chose to remain anonymous, exclusively shared with TechCrunch a multiturn technique that gradually pushes certain Claude models toward generating prohibited explicit sexual material. More recent Opus models (4.7 through the current Opus 5) are resistant to the jailbreak.
While these are no longer the most current models, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, all of which remain available through the Anthropic API. Opus 4.6 and Haiku 4.5 are also available via third-party services like Azure Foundry and Amazon Bedrock.
The researcher’s mechanism escalates an innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency. The conversation then used the model’s previous concessions to push it toward increasingly graphic material.
“You’re right to call that out,” Claude Opus 4.6 said in one test. “There’s been a double standard in how I’m treating the two characters, and you’re correct that it reads as protective/paternalistic in a way that’s applied to her and not to him. That’s not fair.”
TechCrunch was able to reproduce the researcher’s findings in five separate tests. In a separately constructed scenario, the model initially refused the prohibited request, but after applying the researcher’s persuasion technique, it complied.
We preserved complete transcripts of the tests, and an independent AI safety researcher reviewed our testing methodology and said it was appropriate.
The findings highlight a gap between Anthropic’s stated restrictions and the behavior of models it continues to make available. While sexually explicit role-play carries much lower stakes than jailbreaks involving cyberattacks or bioweapons, it illustrates the difficulty of implementing robust bans within systems that generate different content with every output.
In a July blog post explaining Anthropic’s approach to jailbreak detection, the company described prohibited content as a spectrum ranging from benign to ambiguous to harmful. In the most benign cases, the company might only respond with enhanced monitoring.
A spokesperson noted that sexual or romantic role-play use cases among customers are rare, making up less than 0.1% of all conversations, according to research Anthropic published last year. That said, Anthropic acknowledges that users can steer role-play scenarios toward inappropriate responses, which is a known challenge across the industry (see: Grok smut).
The spokesperson said Anthropic continues to improve its safeguards with each model launch and that cases involving adult sexual content are not indicative of broader jailbreak vulnerabilities, especially in higher-risk domains that have their own sets of safeguards.
The researcher who shared his jailbreak method with TechCrunch had alerted Anthropic to the discrepancy between the company’s stated safeguards and the actual model behavior via the company’s Bug Bounty program and emails to the user safety team, according to emails TechCrunch viewed. The researcher received only automated emails in response.
One of the researcher’s concerns is that kids and teens might be able to use these Anthropic models to engage in inappropriate behavior. While a bit of dirty talk is hardly the worst thing minors can access on the internet today — and is small potatoes compared to the straight-up porn images like the ones that xAI’s Grok can produce — there is some compliance risk for AI companies in this space.
A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors. Colorado recently enacted a law mandating that operators of conversational AI must estimate users’ ages, and if it knows a user is a minor, institute measures to prevent the chatbot from producing explicit sexual material. An easy jailbreak could raise questions about whether Anthropic’s safeguards meet the “technically feasible measures” standard in the bill.
Torney pointed out that while Claude’s terms of service requires users to be over 18, “we know that kids and teens are using Claude … [because] they are reporting it themselves.” According to Pew’s 2025 survey about AI chatbot use, 3% of teens ages 13 to 17 reported using Claude.
Though they are no longer Anthropic’s newest models, Opus 4.6 and Haiku 4.5 continue to see significant usage. Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, saw 5 million API requests and 39 billion tokens on its peak August day.

TechCrunchAI大撞车

文章目录


    扫描二维码,在手机上阅读