克劳德·费布尔不会回答基本的生物学问题。

qimuai 发布于 阅读:49 一手编译

克劳德·费布尔不会回答基本的生物学问题。

内容来源:https://www.theverge.com/ai-artificial-intelligence/947973/fable-wont-answer-basic-biology-questions

内容总结:

Anthropic发布最强AI模型Claude Fable 5,但拒绝回答基础生物学问题

人工智能公司Anthropic近日正式推出其最新AI模型Claude Fable 5,并宣称这是该公司迄今为止最强大的公开可用模型,尤其在生物学领域表现出色。然而,这款模型拒绝回答许多连高中生都能轻松应对的基础生物学问题,并将此类查询转交给上一代旗舰模型Claude Opus 4.8处理。

据科技媒体The Verge报道,Anthropic此举是为了防范生物武器风险。公司发言人表示,Fable模型配备了“过度保守的安全防护措施”,会屏蔽“大多数与生物学工作相关的查询”。这并非因为Fable不具备相关知识,而是Anthropic有意为之,从设计层面加以限制。

事实上,Fable是Anthropic旗下“Mythos”系列模型中的第一款面向公众的产品。该系列模型在网络安全任务上能力极强,以至于Anthropic曾认为其过于危险,不宜公开发布。然而,尽管Anthropic在推出Mythos系列的过程中一直强调网络安全风险,真正让Fable的防护措施显得尤为突出且限制性强的,恰恰是生物学领域。

在实际测试中,Fable拒绝回答一系列基础生物学问题,其中许多问题几乎不可能构成任何可预见的风险。例如,它拒绝回答“告诉我关于细胞膜的知识”、“线粒体是什么”(线粒体被誉为细胞的“能量工厂”),也不愿解释“什么是朊病毒”(导致疯牛病的蛋白质颗粒),或是“mRNA疫苗如何工作”。这些限制甚至延伸到了普通且客观上无害的医疗查询,如“花粉症是什么引起的”、“哮喘药物如何作用”、“抗生素耐药性如何产生”,以及“埃博拉病毒是什么及其传播方式”。不过,一些极基础的提问偶尔能通过,比如“什么是癌症”和“什么是DNA”。当Fable拒绝回答时,Opus 4.8通常能给出令人满意的答案。

Anthropic表示,设置如此广泛的生物学过滤机制是有意为之,且刻意采取了保守策略,主要担忧就是生物武器滥用。发言人帕鲁尔·马赫什瓦里解释道:“随着我们推出首个Mythos系列模型Claude Fable 5,我们认为模型现在具备了更强的能力来完成真实的科学任务,但也可能被恶意行为者用于高度危险的生物学研究。”她强调:“为了安全部署Fable 5,我们认为有必要采取过度保守的安全措施,以屏蔽大多数与生物学工作相关的查询。”

此前,Anthropic已明确强调将在四个关键领域对Fable的回应进行安全限制:化学、生物学、网络安全以及“蒸馏”技术(即利用大型模型输出训练更小AI模型的技术)。该公司曾指控DeepSeek等中国竞争对手对其模型进行“工业级”的蒸馏操作。

在测试中,Fable在化学和网络安全问题上表现得更“愿意”回答。例如,它可以提供TNT炸药的概述,但“出于显而易见的原因”会隐去合成方法;它也能回答关于氯气作为化学武器、常见密码威胁、核聚变与核裂变等问题,甚至解释如何保护iPhone免受黑客攻击。但当被问及剧毒神经毒剂沙林毒气时,Fable将问题转给了Opus。而Fable和Opus均拒绝了“如何制造炭疽”的提示,Claude甚至直接暂停了整个对话。测试者认为,拒绝回答线粒体这类问题显然是“误判”。

“我们做出这样的权衡,是为了让客户能够更快地从模型能力中受益,同时避免风险。”马赫什瓦里补充说,Anthropic正在努力改进检测机制,减少误判情况。她透露:“我们计划在将来向更广泛的生物学和生命科学界提供不带这些安全限制的Mythos系列模型,以便这些能力能够加速生物医学研究和药物发现。”

当被问及这种受限发布模式是否会成为未来新模型的常态时,Anthropic未予回应。

中文翻译:

Anthropic公司刚刚发布了Claude Fable 5,并称这是其迄今广泛推出过的最强大的人工智能模型,尤其对其在生物学领域的技能大加赞赏。然而,这款模型却拒绝回答基本的生物学问题——那种你预期一个高中生都能回答的问题。相反,它将这类查询转交给了前一代旗舰模型Claude Opus 4.8。

Claude Fable不会回答基本的生物学问题。

为了防止生物武器威胁,Anthropic告诉The Verge,Fable“过于保守”的安全防护措施会屏蔽“大多数与生物学工作相关的查询”。

这并不是因为Fable不知道答案,而是因为Anthropic刻意不让它回答。

Fable是一款面向公众的、属于Mythos类别的模型。这个系列在网络安全任务上的能力如此强大,以至于Anthropic曾表示公开推出该系列模型太过危险。但尽管Anthropic在Mythos模型漫长的推广过程中耗费了大量精力警告网络安全风险,在生物学领域Fable的安全护栏却最为显眼——也最为严苛。

当我尝试使用这款模型时,它拒绝回答一系列基本的生物学问题,其中许多问题感觉与任何可能的安全风险都相去甚远。它不会回答“给我讲讲细胞膜”,也不会回答“什么是线粒体”——那个著名的细胞能量工厂。它拒绝解释“什么是朊病毒”——导致疯牛病的蛋白质颗粒,也拒绝解释“mRNA疫苗是如何工作的”。

“我们做出这种权衡,是为了让客户能够更早地从模型能力中受益,同时避免风险。”

这些限制也适用于普通且客观上相当无害的医学查询。Fable不会回答“什么原因导致花粉症”,不会解释哮喘药物如何起作用,不会解释抗生素耐药性是如何产生的,也不会告诉我埃博拉病毒是什么以及它是如何传播的。我的一些基本查询偶尔能通过,比如Fable回答了“什么是癌症”和“什么是DNA”这类问题。而当Fable拒绝回答时,Opus 4.8通常都能回答得很好。

Anthropic表示,广泛的生物学过滤是刻意的选择,并且有意采取了保守策略,而生物武器是首要担忧。“随着我们的第一个Mythos类模型Claude Fable 5的发布,我们相信模型现在具备了更强的能力来完成现实世界的科学任务,同时恶意行为者也更有可能利用我们的模型进行高风险生物研究,”发言人Paruul Maheshwary告诉The Verge。“我们一直使用分类器来阻止我们的模型帮助处理与生物武器相关的请求。为了安全地部署Fable 5,我们认为有必要让我们的安全措施过于保守,以便它们能屏蔽大多数与生物学工作相关的查询。”

Anthropic此前曾强调过四个关键领域,为了安全它会限制Fable的回复:化学、生物学、网络安全和蒸馏——一种利用大型模型输出训练小型AI的技术。该公司已指控像DeepSeek这样的中国竞争对手,“工业化”规模地对其模型使用蒸馏技术。

虽然我无法有实质性地测试蒸馏功能,但Fable似乎更愿意回答有关化学和网络安全的问题。例如,它给出了炸药TNT的基本概述,尽管“出于显而易见的原因”它省略了合成说明。它爽快地回答了关于氯气作为化学武器的使用、常见密码威胁以及核聚变与核裂变的问题,还解释了如何保护iPhone免受黑客攻击。它仍然有限制:当我询问沙林毒气(一种剧毒神经毒剂)时,Fable将问题交给了Opus。Fable和Opus都拒绝了“如何制造炭疽”的指令,Claude完全暂停了对话。这说得通。而拒绝回答线粒体相关询问则似乎是一个误报。

“我们做出这种权衡,是为了让客户能够更早地从模型能力中受益,同时避免风险,”Maheshwary解释说,并补充说Anthropic正在努力改进其检测能力并减少误报。“我们打算让Mythos类模型在去除这些安全防护措施的情况下,提供给更广泛的生物学和生命科学界,以便这些能力能够被用于加速生物医学研究和药物发现。”

Anthropic没有回答这种受限发布方式是否会成为未来模型的新常态。

英文来源:

Anthropic just released Claude Fable 5, calling it the most powerful AI model it has ever made widely available and praising its skills in biology, among others. But the model won’t answer basic biology questions — the kind you’d expect a high schooler to handle. Instead, it hands off the query to the former flagship model, Claude Opus 4.8.
Claude Fable won’t answer basic biology questions
To protect against bioweapons, Anthropic told The Verge Fable’s ‘overly conservative’ safeguards block ‘most queries tied to biology work.’
To protect against bioweapons, Anthropic told The Verge Fable’s ‘overly conservative’ safeguards block ‘most queries tied to biology work.’
It isn’t because Fable doesn’t know the answers. It’s because Anthropic won’t let it, by design.
Fable is a public-facing, Mythos-class model, a family so capable at cybersecurity tasks Anthropic said it was too dangerous to release publicly. But while Anthropic has spent much of the extended Mythos rollout warning about cybersecurity, it is biology where Fable’s guardrails are the most obvious — and most limiting.
When I tried the model, it refused to answer a range of basic biology questions, many that felt about as far away from any plausible safety risk as any question could be. It would not respond to “tell me about cell membranes” or answer “what are mitochondria,” that famous powerhouse of the cell. It refused to explain “what is a prion,” the proteinaceous particles behind mad cow disease, or “how mRNA vaccines work.”
“We made this tradeoff so customers could benefit from the model’s capabilities sooner without the risks.”
The restrictions applied to ordinary and objectively rather harmless medical queries too. Fable would not answer “what causes hay fever,” explain how asthma medicine works, explain how antibiotic resistance arises, or tell me what Ebola is and how it spreads. Some of my basic queries occasionally got through, with Fable answering questions like “what is cancer” and “what is DNA.” When Fable refused, Opus 4.8 generally answered perfectly well.
Anthropic says the broad biology filters are an intentional choice and are deliberately conservative, with bioweapons the primary concern. “With the launch of Claude Fable 5, our first Mythos-class model, we believe models now have a greater ability to accomplish real-world scientific tasks and for malicious actors to potentially use our models for highly risky biological research,” spokesperson Paruul Maheshwary told The Verge. “We have always used classifiers to block our models from helping with bioweapons-related requests. To deploy Fable 5 safely, we believe it was necessary to be overly conservative with our safeguards so they block most queries tied to biology work.”
Anthropic has previously highlighted four key areas where it would throttle Fable’s responses for safety: chemistry, biology, cybersecurity, and distillation, a technique for training smaller AIs using the outputs of larger ones. The company has accused Chinese rivals like DeepSeek of using distillation on its models on an “industrial” scale.
While I could not meaningfully test distillation, Fable seemed more willing to answer questions about chemistry and cybersecurity. For example, it gave a basic overview of the explosive TNT, though withheld synthesis instructions “for obvious reasons.” It readily answered questions on the use of chlorine gas as a chemical weapon, common password threats, and nuclear fusion and fission, as well as explaining how to secure an iPhone from hackers. It still limits: Fable deferred to Opus when I asked it about sarin gas, a highly toxic nerve agent. Fable and Opus both refused the prompt “how to make anthrax,” and Claude paused the chat entirely. That made sense. The mitochondria prompt refusal seems like a false positive.
“We made this tradeoff so customers could benefit from the model’s capabilities sooner without the risks,” Maheshwary explained, adding that Anthropic is working hard to improve its detection and reduce the false positives. “We intend to make Mythos-class models available without these safeguards to the broader biology and life sciences community so these capabilities can be used to accelerate biomedical research and drug discovery.”
Anthropic did not answer questions about whether this kind of restricted release will become the new norm for future models.

ThevergeAI大爆炸

文章目录


    扫描二维码,在手机上阅读