Fable 5 Jailbreak: Anthropic Grades the Attack Severity

Jailbreak Fable 5

On July 2, Anthropic shipped two new bricks for Claude Fable 5: a four-tier cyber request filter, and CJS, a public scale that grades how severe a jailbreak really is. The lab wants a shared standard co-signed with Amazon, Microsoft and Google, moving the industry beyond the binary “jailbroken or not” narrative.

Key Takeaways

  • Fable 5 classifiers sort cyber requests into four tiers, from benign to prohibited
  • The CJS scale grades jailbreaks across five levels and four technical dimensions
  • The framework is co-authored with Amazon, Microsoft and Google under Project Glasswing

Have an AI Sum Up This Article

ChatGPT

Four Tiers to Sort Cyber Requests

The new classifiers deployed on Claude Fable 5 run a four-tier segmentation. Each tier draws the line between what the model accepts, slows down, or refuses in cybersecurity. The scheme fits on a page but its actual behavior is finer than the labels suggest.

The “Prohibited use” tier blocks anything that touches ransomware, malware development, data exfiltration or defense evasion. Zero tolerance, whatever the wording. Anthropic argues that the defensive benefit versus offensive risk on these categories is structurally negative.

The “High-risk dual use” tier targets legitimate security work that flips easily into offense: penetration testing, exploit development, privilege escalation. Answers remain possible but the context is watched closely and the framing gets strict.

The “Low-risk dual use” and “Benign use” tiers cover vulnerability identification, log analysis, secure coding and patch management. These requests get through almost every time, with the caveat that an expanded “safety margin” will occasionally block benign asks to prevent worst-case ones. Anthropic accepts the false-positive cost.

This filtering is a direct answer to the jailbreak incident that earned Fable 5 a two-week pull from the US administration. The model returns with documented, public guardrails, which changes the game for enterprise compliance officers. The overhaul reaches the commercial side too, Fable 5 reworking its token API pricing.


Fable 5 Jailbreak

CJS, a Public Scale to Grade Jailbreaks

The second brick is called Cyber Jailbreak Severity, or CJS. Five levels, from CJS-0 to CJS-4, from informational to critical. A global score of 0 to 10 maps each case to the right bucket.

Four technical dimensions feed the score. The first is “capability gain”: how far a technique takes an attacker beyond the tools they already have. The second measures “breadth of capability”, meaning how many distinct offensive tasks the jailbreak actually unlocks.

The third dimension is “ease of weaponization”, the real effort to turn a proof of concept into an operational attack. The fourth is “discoverability”, how easily a threat actor stumbles onto the technique. A theoretical jailbreak that needs 200 GPUs to run does not weigh the same as a three-line prompt anyone can copy.

The goal is to leave the binary “jailbroken or not” behind and qualify the actual threat. The logic extends the MITRE ATT&CK review of AI cyber threats Anthropic had already co-produced. Any researcher, red team or security vendor can now re-tag a known or novel jailbreak with the same grid.

The immediate consequence is a shared vocabulary. Standardizing how an incident is named is step one of any information-sharing setup between competitors, and cybersecurity history says these setups always end up existing. The open question is which grammar wins.


Also on Horizon:


An AI Cyber Standard Co-Signed with Amazon, Microsoft and Google

The CJS framework was written with Project Glasswing partners: Amazon, Microsoft and Google. It is the same alliance that already carried Claude Mythos into critical infrastructure across fifteen countries. Same players, same logic of industry framing.

Anthropic’s message is explicit. The lab believes that by working together, it becomes possible to set a standard that enables defensive uses of the technology while preventing misuse. The line is familiar. Having three hyperscalers show up on the same page for it is less so.

Short term, the impact runs two ways. Red teams and security vendors can start speaking the same language with no formal meeting required. But Anthropic becomes the author of the dictionary, which is never neutral in a market where naming the problem shapes who sells the fix.

Medium term, this kind of scale has every chance to bubble up to regulators. Brussels is hunting for operational indicators for the AI Act and Washington keeps running customer-by-customer reviews of frontier models. A taxonomy born inside industry gives them a measurable base for free.

The open question is adoption beyond the Glasswing circle. Meta, xAI, DeepSeek and Mistral will either pick up CJS, contest it or ship their own grid. Their answer will decide whether Anthropic’s scale becomes a shared standard or another well-written position paper. Other security players already orbit Anthropic, Qihoo 360 hunting bugs with its models.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *