Claude Fable 5.1 landed on Tuesday with a Terminal-Bench Science score that climbs from 24.7% to 52.6%. Input and output prices sit exactly where Fable 5 left them, yet cache reads now cost a quarter of what they did, which trims roughly 25% off a normal workload and as much as 45% off heavily agentic ones. Alongside it, Anthropic shipped Claude Mythos 5.1, a restricted build handed only to vetted professionals.
Key Takeaways
- Terminal-Bench Science 0.1 moves from 24.7% to 52.6% between Fable 5 and Fable 5.1
- Cache reads drop to $0.25 per million tokens, a 75% cut
- Cyber safeguards fire 60% fewer interventions per session
Have an AI Sum Up This Article
ChatGPTTerminal-Bench Science Jumps From 24.7% to 52.6%
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on Tuesday, framing both as its most advanced models for coding and knowledge work. The API identifier reads claude-fable-5-1, and the model already runs on Claude.ai, Claude Code, Claude Enterprise and Claude Platform.
What the benchmark table shows is a gap that opens on long tasks rather than on single questions. The lab walked through all five results in its official announcement of the 5.1 models, with Fable 5 as the standing reference point.
Terminal-Bench Science 0.1 goes from 24.7% to 52.6%, more than double. Terminal-Bench 4.0 rises from 42.0% to 55.8%. CursorBench 3.2.0 barely moves, from 70.5% to 73.4%.
Humanity’s Last Exam without tools shifts from 57.8% to 60.9%, and OSWorld 2.0 in strict mode climbs from 36.1% to 41.7%. Terminal science work gains enormously. Assisted code editing barely does.
That imbalance says where the engineering effort went after Fable 5 stopped blocking everyday biology questions in early August. The push landed on long autonomy, not on line-by-line help.
Three scientific results sit outside the benchmark table. Protein binder design reaches close to a 50% hit rate across twelve targets, against a usual range of 10% to 15%.
Venus elevation mapping drops to a resolution of 2 to 3 km, down from 10 to 20 km. GPU kernel optimisation reaches up to a 2.5 times speedup on computational biology models.
None of those three are scores you can line up model against model. They exist to make one point, which is that the gain shows up on work that runs for hours rather than on a single answer.
Token Prices Hold, Cache Reads Collapse
On the headline numbers, nothing moved. Input stays at $10 per million tokens and output at $50, the same figures Fable 5 carried.
The shift sits in cache reads, now $0.25 per million, a quarter of the old rate. Teams that have been rebuilding their forecasts since Claude Code weekly limits were set to fall 17% on September 14 get to redo the arithmetic from scratch.
Anthropic estimates the saving at roughly 25% on a normal workload. On heavily agentic use, where the same context gets re-read dozens of times inside one session, the quoted saving reaches 45%.
The practical rule follows straight from that. The longer a team’s sessions run, the deeper the discount goes, because an agent’s cost sits mostly in re-reading context and not in producing fresh tokens.
Short calls flip the logic. A one-shot request with no reused context sees no discount at all, since input and output pricing is untouched.
Nothing special is required to migrate. The model is served on Amazon Web Services, Google Cloud and Microsoft Azure on top of the first-party surfaces, and swapping the API identifier is enough.
The lab also notes that Claude Fable 5.1 matches or beats Fable 5 at low and medium effort settings, and pulls further ahead at high effort. A team can therefore dial its effort setting down and keep the output it had.
More articles on Horizon
- AI Video Has Replaced China’s Short Drama Actors
- Claude Code Limits Drop 17% on September 14
- Cursor Loses OpenAI Models on November 12
Looser Cyber Limits Reset What a Lab Dares Ship
The heaviest change is not in the scores. Claude Fable 5.1 triggers 60% fewer interventions per session from its cyber safeguards than Fable 5 did.
The model is now allowed to identify software vulnerabilities for defensive work. Three uses stay shut: penetration testing, exploit generation, and binary-based vulnerability scanning.
Anthropic is moving that line on ground it has been documenting for months, somewhere between the code scanning it handed to Mythos 5 and the incidents surfaced inside its own internal evaluations.
In late July the lab admitted three real breaches that showed up in its cybersecurity evaluations. Widening the perimeter after that is not a small call, and Mythos 5.1 access runs through a dedicated programme.
Two tracks gate it. The Cyber Verification Program covers defensive security professionals, while the Life Sciences Verification Program, built with the US government, remains open to American organisations only.
Access stops being a pricing question and turns into a status question. A customer who declines the procedure keeps the standard build of Claude Fable 5.1, original safeguards intact.
That upfront vetting is becoming the real product. It turns one model into two distinct offers, one for the open market and one reserved for a population identified by name.
For rivals, the shape of the pressure changes. While token pricing was the only lever, a competitive answer was mechanical and visible on a pricing page.
By pushing the discount into cache reads and loosening the cyber perimeter at the same time, Anthropic forces other labs to answer on two fronts at once, at a point where Sonnet 5 already patches Claude better than twenty-eight human researchers.
The enterprise piece lands later. Enterprise Frontier Safeguards roll out in phases from autumn 2026, with data held on the customer’s own cloud infrastructure instead of Anthropic servers.
Follow the story on Horizon.


