Fable 5 Releases Safety Classifier Details and Jailbreak Severity Framework
Fable 5 is now available globally, and the development team has disclosed additional information regarding its security features and approach to preventing misuse. The announcement focuses on two key areas: the safety classifiers built into the model and a proposed framework for evaluating AI jailbreak severity.
Safety Classifiers and Cybersecurity Safeguards
The model includes AI-based safety classifiers designed to detect and block potentially dangerous cybersecurity activities. Rather than preventing all cybersecurity-related use, the team has implemented a targeted approach that distinguishes between different types of requests. This recognizes that many cybersecurity capabilities serve dual purposes—they can be used defensively by security professionals or maliciously by attackers.
The classifiers operate within four categories:
- Prohibited use: Activities with minimal defensive benefit and high harm potential, such as ransomware creation, cyberattacks on critical infrastructure, and anti-forensics techniques
- High-risk dual use: Capabilities widely exploited by malicious actors but with legitimate defensive applications
- Low-risk dual use: Primarily defensive activities that could also assist attackers
- Benign use: Activities without harmful applications
The classifiers employ a deliberate "safety margin" that blocks some legitimate requests to prevent harmful ones. For Fable 5, this margin is larger than in previous iterations to provide additional protection.
Jailbreak Severity Framework
The team has developed an early version of a jailbreak severity framework in collaboration with Glasswing partners. Jailbreaks are techniques that bypass AI safeguards through unconventional prompting. Currently, no standardized system exists for rating their severity, which hampers communication between developers, governments, and policymakers.
The framework aims to establish consistent terminology for describing how dangerous a particular jailbreak is. The team has invited feedback on this framework and encourages security researchers to submit discovered vulnerabilities through a new HackerOne program.
The company emphasizes that broader safeguards—including access controls, model training, and monitoring—work alongside classifiers to create multiple protective layers.
More from Technology & AI
Anthropic Launches Claude Opus 5 AI Model with Advanced Performance
26 July, 2026 · 09:12
London Stock Exchange Plans 24/5 Trading Venue for Digital Markets
23 July, 2026 · 14:27
Anthropic Launches Economic Index Connector for Claude AI
23 July, 2026 · 09:55
Anthropic Launches Initiative to Address Public Concerns About AI
22 July, 2026 · 09:07