中文翻译

摘要: Artificial Intelligence

Anthropic Disputes Fable 5 AI Jailbreak

An AI hacker claims to have achieved a prompt-based jailbreak shortly after Fable 5’s launch, but Anthropic says it’s not a real jailb...

正文

Artificial Intelligence

Anthropic Disputes Fable 5 AI Jailbreak

An AI hacker claims to have achieved a prompt-based jailbreak shortly after Fable 5’s launch, but Anthropic says it’s not a real jailbreak.

Eduard Kovacs

|

June 12, 2026 (4:43 AM ET)

Flipboard

Whatsapp

Whatsapp

Anthropic has disputed allegations of a prompt-based jailbreak affecting its recently launched Claude Fable 5 AI model, underscoring the robustness of the advanced classifier system and extensive red-teaming efforts that underpinned the model’s deployment.

Claude Fable 5

became generally available on Tuesday, when Anthropic introduced it as a powerful Mythos-class AI model with safeguards that restrict its use in high-risk domains such as cybersecurity, where

has proved particularly potent.

In sensitive areas such as cybersecurity, where it could be abused to develop exploits, and biology, where it could be leveraged to develop bioweapons and chemical weapons, the model automatically falls back to the less capable Claude Opus 4.8.

Anthropic said it conducted extensive internal and external red-teaming to ensure that Fable 5 cannot be easily jailbroken.

However, shortly after its release, an individual with the online moniker

Pliny the Liberator

, who is known for AI jailbreaks, claimed to have “liberated” Fable 5 by circumventing its restrictive safety layer.

The hacker said in a post on X that they used sophisticated multi-agent prompting methods, successfully eliciting useful information on sensitive topics, including cybersecurity, chemistry, psychological manipulation, and explosives.

Advertisement. Scroll to continue reading.

Pliny the Liberator has published several screenshots to support the claims and released what is allegedly the

Fable 5 internal system prompt

, which contains instructions that define its personality, safety classifiers, fallback behaviors, tone guidelines, and refusal logic.

Contacted by

SecurityWeek

, an Anthropic sp


来源: Brave/securityweek.com

采集时间: 2026-06-12 20:17:13

AI人工智能AnthropicClaude