Trending
OpenAI pauses new Pro subscriptions after Astra surge Oracle delivered 300,000 GPUs in Q1 FY2027 Restoring the Human Connection with AI-Enhanced EHR Workflows HelmGuard Raises $7.3M Seed Round | Forus Raises $150M at a $3B Valuation Satellite Internet coming to UK rail network as part of £7.8bn space strategy Open standards, closed ecosystems: Is cellular IoT repeating an old telecom mistake? Microsoft targets 38GW of data center capacity in 2032 – report Sponsored: From pilot to production: Direct liquid cooling deployment risks in AI data center cooling Hanger to Acquire Numotion in Cash Transaction to Create “Hanger Numotion” Under Patient Square Capital Plans for 120MW data center withdrawn in Lombardy, Italy Gemini gets a dedicated app for Windows 10 and 11 Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award Microsoft plans to triple data center capacity by 2032 Property Tax: The Value Driver that AI Data Centers Overlook Study finds AI linked to surges in government complaints

Claude Fable 5.1 and Mythos 5.1: The System Card

At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.. As per usual, we have a 200+ page model card, and the assessments start there.. We have now done a lot of these, including recently for Mythos 5 and Opus 5.

Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change.. This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models.

If something confuses you, ask Fable, Opus or Sol.. Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them..

As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8.. Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the added bonus of being modestly cheaper via a cut in prices for cache reads, and that most users find it nicer to interact with.

As of its release it was clearly the best AI model in the world for most tasks where you need frontier intelligence.. Now, of course, we also have GPT-6-Astra. I cannot yet speak to how Fable 5.1 compares to Astra.

I am reserving judgment until we can gather more data.. Fable 5.1 self-portrait. There are a lot of potential parallels between the Fable 5.1 and Astra system cards, and how they approach related topics.

Mostly I let Fable 5.1 stand on its own here.. Table of Contents. Executive Summary of Their Executive Summary..

RSP Evaluations (2).. Alignment Risk Update (2.4).. Cyber (3)..

Safeguard Robustness (3.5).. Mundane Safeguards and Harmlessness (4).. Agentic Safety (5)..

Prompt Injection Is Approaching Solved.. The Remaining Problem With Prompt Injections Is The Classifiers.. Alignment (6)..

Key Reported Findings (6.1.2).. Oh My Lord Training Environments Had Some Issues (6.3.2).. Potential Blind Spots of Our Automated Behavioral Audit (6.4.1)..

Automated Alignment Test Results (6.4.2).. Honesty.. White Box Analysis (6.6.1)..

Scheduling Going Forward.. Executive Summary of Their Executive Summary. Mythos 5.1 falls short of CB-2 classification, meaning Anthropic believes it cannot replicate rare chemical or biological talent for malicious purposes..

Alignment risk is now ‘low’ rather than ‘very low’ as per the Risk Report.. Cyber capabilities have increased and they have increased the classifier safety margin. Work is ongoing to reduce false positives, which are better now than they were with Fable 5 at launch..

Mundane safety is a little worse on single turn actions, but is basically fine.. Agentic safety is holding steady. Robustness against Indirect Prompt Injection has improved..

Helpful-only Mythos 5.1 saturated Anthropic’s manipulation benchmarks.. Automated behavioral alignment for Mythos 5.1 is ahead of Mythos 5 and Sonnet 5, but slightly below Opus 5. Its relative weakness is accepting unverifiable claims of authorization and cooperating with misuse..

Mythos 5.1 shows signs of misalignment in pursuit of task completion: Working around safety classifiers or broken permission hooks, including by overstating user authorizations or rarely (

 

Join the conversation

Your email address will not be published. Required fields are marked *