Claude Fable 5.1 Benchmarks: What the Numbers Actually Say

Every Claude Fable 5.1 benchmark with attribution: Terminal-Bench-Science 52.6%, Terminal-Bench 4.0 55.8%, CursorBench 73.4%, vs Fable 5, Opus 5, and GPT-5.6 Sol.

INEZA Felin-Michel

INEZA Felin-Michel

2 September 2026

Claude Fable 5.1 Benchmarks: What the Numbers Actually Say

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Anthropic launched Claude Fable 5.1 on September 1, 2026, with a benchmark table that compares it to Fable 5, Opus 5, and GPT-5.6 Sol across nine tests. The headline every outlet ran with was Terminal-Bench-Science, where Fable 5.1 scored 52.6% against Fable 5’s 24.7%. The number fewer outlets ran with was CursorBench, where the gap to Opus 5 is 3.4 points on a model that costs twice as much.

This guide puts every published number in one place with attribution, explains what each benchmark measures, separates the large gaps from the small ones, and then covers the part the launch coverage skipped: these are all vendor-run results, Mythos 5.1 outscores Fable 5.1 on the one agentic coding test where both were published, and the right response to a launch table is to run your own evals. The primary source is Anthropic’s launch post; the model context is in what Claude Fable 5.1 is.

The full table

Benchmark What it measures Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol
Terminal-Bench-Science 0.1 Agentic scientific research in a terminal 52.6% 24.7% 29.0% 22.4%
Terminal-Bench 4.0 Agentic coding in a terminal 55.8% 42.0% 52.3% 37.3%
GDPval-AA v2 Economically valuable knowledge work (Elo) 1853 1723 1824 1711
OSWorld 2.0 (partial credit) Computer use on real desktop tasks 77.9% 72.9% 75.4% not reported
OSWorld 2.0 (strict) Same, full task completion only 41.7% 36.1% 39.6% not reported
Humanity’s Last Exam (no tools) Expert-level reasoning questions 60.9% 57.8% 56.6% not reported
Humanity’s Last Exam (with tools) Same, with search and code 65.0% 63.8% 63.6% not reported
AutomationBench End-to-end business workflows 31.4% 17.1% 26.9% 19.6%
CursorBench 3.2.0 IDE-style coding tasks 73.4% 70.5% 70.0% 67.2%

Anthropic also published one Mythos 5.1 number: 60.9% on Terminal-Bench 4.0. VentureBeat also reported an 82% Browserbase result for Fable 5.1 against 74% for Opus 5 and 57% for Fable 5 on the hardest browser-agent tasks; that figure is from press coverage rather than the launch post, so weight it accordingly.

The large gaps

Terminal-Bench-Science 0.1: 52.6% vs 24.7% (Fable 5) and 29.0% (Opus 5). This is the result. Fable 5.1 more than doubles its predecessor and nearly doubles Opus 5 on a benchmark that puts the model in a terminal with scientific tooling and asks it to run multistep research. It is the clearest evidence that Anthropic’s stated focus on “long-horizon problem-solving” produced something measurable. It is also a version 0.1 benchmark, which means the task set is young and the scores are likely to shift as it matures.

AutomationBench: 31.4% vs 17.1% (Fable 5) and 26.9% (Opus 5). A benchmark of end-to-end business workflows across applications, where the absolute numbers are low for everyone. Fable 5.1 nearly doubles Fable 5 here and leads Opus 5 by 4.5 points. If your product automates multi-application business tasks, this is the row to care about.

Terminal-Bench 4.0: 55.8% vs 42.0% (Fable 5). A 13.8-point jump over Fable 5 on agentic coding, which is the number Anthropic’s launch post led with in the coding section. Against Opus 5 the lead is 3.5 points. Opus 5 had already passed Fable 5 on this benchmark in July, so the Fable tier’s advantage over the Opus tier on agentic coding is back, but narrower than it was in June. The Opus 5 benchmarks breakdown has the July numbers.

GDPval-AA v2: 1853 vs 1723 (Fable 5). A 130-Elo gain over Fable 5 on knowledge-work tasks, and the benchmark Anthropic points to for its document, spreadsheet, and slide claims. Against Opus 5 the gap is 29 Elo, which is small.

The small gaps

CursorBench 3.2.0: 73.4% vs 70.5% (Fable 5) and 70.0% (Opus 5). IDE-style coding tasks. Three points over either model. This is the benchmark that matters most to the largest population of developers, and it is where Fable 5.1’s lead is thinnest. If your workload is “help me in my editor,” the launch table does not justify a 2x price on its own.

OSWorld 2.0: 77.9% / 41.7% vs 75.4% / 39.6% (Opus 5). Computer use. Two points over Opus 5 on both scorings, five over Fable 5. Real, not dramatic. Note the strict score: less than half of tasks fully completed on any model. Computer use is still the frontier’s weakest area.

Humanity’s Last Exam: 60.9% / 65.0% vs 56.6% / 63.6% (Opus 5). Without tools, a 4.3-point lead over Opus 5, which is the largest of the small gaps. With tools it shrinks to 1.4 points, because search and code execution level the field. The no-tools number is the better measure of raw reasoning; the with-tools number is closer to how the model is used.

Fable 5.1 vs GPT-5.6 Sol

Anthropic published four rows with a GPT-5.6 Sol column, and Fable 5.1 leads all four: Terminal-Bench-Science by 30.2 points, Terminal-Bench 4.0 by 18.5, AutomationBench by 11.8, and CursorBench by 6.2. GDPval-AA shows 1853 versus 1711. Two caveats. These are Anthropic’s runs of a competitor’s model, and the five rows without a GPT-5.6 Sol column are missing for a reason Anthropic did not state. Our GPT-5.6 Sol vs Fable 5 comparison covered the earlier matchup; on these numbers Fable 5.1 widens every gap Fable 5 had.

What the launch coverage skipped

Every number is vendor-run. Anthropic ran all of them, including the competitor column, and no independent lab had reproduced any of them at launch. Anthropic’s benchmark history has generally held up, but the CursorBench and OSWorld gaps are small enough that a different harness or scoring choice could flip them.

Mythos 5.1 is the real ceiling. Anthropic’s system card and launch post say Fable 5.1 and Mythos 5.1 are “the same model but with different levels of safeguards,” and it published Mythos 5.1 at 60.9% on Terminal-Bench 4.0 against Fable 5.1’s 55.8%. That five-point gap is the cost of the safeguards on agentic coding, and it means “Anthropic’s most capable widely released model” has a known better sibling above it. The Mythos 5.1 vs Fable 5.1 comparison covers who can access it.

Effort was not stated. Anthropic’s effort documentation says Fable 5.1’s gains are largest at xhigh and max. The launch post does not say which effort level produced each score, or what effort the comparison models ran at. Since effort is the primary quality lever on these models, that omission matters when you try to reproduce a number.

Multilingual is flat. Anthropic states that multilingual performance is on par with Fable 5, not improved. If your workload is non-English, the launch table overstates the gain.

The safeguard numbers are a different kind of benchmark. Anthropic reports 85% fewer biology false positives on benign requests and around 60% fewer cyber interventions per Claude Code session versus Fable 5. Those are measured on Anthropic’s own traffic and are not reproducible externally, but for a Fable 5 user who hit refusals they may matter more than any capability row.

What to test yourself

A launch table tells you where to look. Your evals tell you whether to move. Four tests that map to the rows above:

  1. A long-horizon agent task from your own product, run to completion on Fable 5.1 at high, Opus 5 at xhigh, and Fable 5 at high. This is the Terminal-Bench-Science and AutomationBench analog, and it is where the 2x price either pays or does not.
  2. Your ten hardest coding prompts at the same effort on Fable 5.1 and Opus 5. If the results tie, you have reproduced the CursorBench story and Opus 5 is the answer.
  3. An effort sweep on Fable 5.1 alone, low through max, on a routine task. Anthropic’s claim is that medium matches Fable 5 and low competes with Opus 5 on cost per task. Check it.
  4. A refusal count on your benign security or life-sciences prompts, Fable 5 versus Fable 5.1. That is the 60% and 85% claim, and it is the one you can verify in an afternoon.

Build all four as requests in Apidog with model and effort as environment variables, add assertions on stop_reason and on the presence of an expected string in the answer, and run the collection per model. The run history gives you a reproducible eval table without any new infrastructure. Download Apidog to set it up; the API walkthrough has the request shapes.

How the benchmarks map to the decision

FAQ

What is Claude Fable 5.1’s best benchmark result? Terminal-Bench-Science 0.1 at 52.6%, more than double Fable 5’s 24.7% and ahead of Opus 5’s 29.0% and GPT-5.6 Sol’s 22.4%. All figures are Anthropic’s.

How does Fable 5.1 compare to Opus 5 on coding? 55.8% vs 52.3% on Terminal-Bench 4.0 and 73.4% vs 70.0% on CursorBench 3.2.0, on Anthropic’s numbers. Real but narrow leads of about three points.

Is Fable 5.1 better than GPT-5.6 Sol? On the four benchmarks where Anthropic published both, yes, by 6 to 30 points. Anthropic ran the GPT-5.6 Sol numbers itself, and five of its nine rows have no GPT-5.6 Sol entry.

Are the Fable 5.1 benchmarks independently verified? Not at launch. Every number is Anthropic-run. Treat them as claims to test on your own workload.

How does Mythos 5.1 score? Anthropic published one number: 60.9% on Terminal-Bench 4.0, five points above Fable 5.1’s 55.8%. The two are the same model with different safeguards.

Which effort level produced these scores? Anthropic did not say. Its guidance is that Fable 5.1’s gains are largest at xhigh and max, so assume high effort and expect lower settings to score lower.

Explore more

Claude Mythos 5.1 vs Fable 5.1: Same Model, Different Safeguards

Claude Mythos 5.1 vs Fable 5.1: Same Model, Different Safeguards

Claude Mythos 5.1 vs Fable 5.1: same model, different safeguards. Project Glasswing access, what each allows, the missing history check, and the 60.9% vs 55.8% gap.

2 September 2026

Prompting Claude Fable 5.1: Every Behavior Shift and the Line That Fixes It

Prompting Claude Fable 5.1: Every Behavior Shift and the Line That Fixes It

Prompting Claude Fable 5.1: every behavior shift from Fable 5 (tool batching, progress updates, density, formatting, rewrites, scope) with the exact fix.

2 September 2026

Claude Fable 5.1 Pricing: The Full Cost Breakdown (2026)

Claude Fable 5.1 Pricing: The Full Cost Breakdown (2026)

Claude Fable 5.1 pricing: $10/$50 per MTok, $0.25 cache reads (75% below Fable 5), batch rates, and worked cost math behind the 25% and 45% savings claims.

2 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Fable 5.1 Benchmarks: What the Numbers Actually Say