AIThis post was created with the assistance of artificial intelligence (AI).

By Thorsten Meyer

Artificial Analysis has Claude Fable 5.1 at the top of its Intelligence Index — 66 at max effort, the highest score the benchmark has ever recorded, ahead of Claude Opus 5 at 63, GPT-5.6 Sol and Grok 4.6 at 61, and a field of nearly two hundred models behind them. That’s a real result and I’ll give it its due. But I apply the same discipline to Anthropic’s model that I apply to everyone’s, so here’s the sentence the headline leaves out: Fable 5.1 also costs about 20% more per task than the model it replaces, because it’s more verbose. “Smartest on the index” and “cheapest per task” are different claims, and the interesting analysis lives in the gap between them. One disclosure up front, because it’s the honest thing to note and Artificial Analysis notes it themselves: they supported Anthropic with pre-release evaluation of this model. It doesn’t invalidate the numbers — AA is a credible independent benchmarker — but it’s the kind of relationship I’d flag for any vendor, so I’m flagging it here.

Let me take the win, then the cost line, then the asterisks that keep the win honest.

AI DISPATCH · REALITY CHECKClaude Fable 5.1 · AA Intelligence Index · 29 Aug 2026
“Smartest on the index” ≠ “cheapest per task”
Fable 5.1 Tops the Index — Now Read the Cost Line

A real new high on Artificial Analysis’s Index (66, above Opus 5’s 63) — and about 20% more per task than Fable 5, because it’s verbose. The interesting analysis lives in that gap.

66 (max)
AA Index · highest measured
$3.76/task
Max · ~20% > Fable 5 · 1.6× Opus 5
~1.7×
Output tokens vs Fable 5 (verbose)
−75%
Cache read cut · $1 → $0.25 / 1M
The knob that decides your budget — effort level, not the headline 66
low
58 · $0.77
xhigh
65 · $2.72
max
66 · $3.76
5 effort levels span 11× in tokens (58→66). The crown (66) is the least economical corner. xhigh scores 65 at $2.72 — still beats Opus 5 (63, $2.34) at a smaller premium than max. Most deployments want a notch down.
The cache cut helps — but only some workloads
Cache-heavy agentic → you save
Long tool-using sessions read the same context repeatedly. The 75% cut saves ~$1.40/task; ~25–45% lower overall. Without it, Fable 5.1 would cost ~$5.16/task.
Novel reasoning → you pay
Fresh output tokens aren’t cached, so the cut barely touches you — you just eat the ~20% verbosity premium. Same model, opposite cost outcome. Your token mix decides.
The asterisks that keep the win honest
~“Tops the leaderboard” is sometimes within the noise. On agentic work its leads over Opus 5 are within the confidence interval or effectively tied — ahead on analysis, behind on presentation.
!Record accuracy (67.2%) comes with more hallucination. It attempts more questions (93.4%), so it gets more right and more wrong than its predecessor.
iYou’re measuring the model + its safety fallback (~4% of output tokens routed to Opus 4.8/5). And AA disclosed it supported Anthropic with pre-release evaluation.

The win is real

The gains are broad and third-party-measured, which matters more than any single number. Fable 5.1 adds four points on the Intelligence Index over Fable 5, and the composite spans reasoning, coding, knowledge, and math rather than one cherry-picked task. On Humanity's Last Exam it scores 59.1%, up from Fable 5's 55.5%. It posts the narrowly-highest scores Artificial Analysis has measured on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), gains nine points on a banking agent benchmark, and on the two agentic knowledge-work evaluations it sets the highest Elo scores AA has recorded. That these come from an outside evaluator running a fixed suite, rather than from Anthropic's own slides, is the part that earns the result credibility. This is a genuine frontier step, not a benchmark stunt.

Amazon

AI language model API subscription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Now the cost line

Here's what the leaderboard crown hides. At max effort, Fable 5.1 costs about $3.76 per Intelligence Index task — roughly 20% more than Fable 5 at $3.14, and about 1.6 times Claude Opus 5 at $2.34. The reason isn't the per-token price, which is unchanged at $10 per million input and $50 per million output. The reason is that Fable 5.1 is talkative: it generates around 1.7 times the output tokens of Fable 5, and on the Index it burned about 140 million output tokens against a median of 71 million for comparable models. You're paying for a model that thinks out loud at length, and on the frontier, output tokens are where the bill accumulates.

Anthropic clearly saw this coming, because alongside the model it cut the price of cache reads by 75% — from $1 to $0.25 per million cached input tokens — while leaving everything else unchanged. That's a genuinely material move, and it's aimed precisely at the workloads where it helps most. In agentic work, the large majority of input tokens are cache reads — the same context read over and over across a long tool-using session — so the cut saves around $1.40 per task there; without it, Fable 5.1 would run about $5.16 per task. The honest translation: if your workload is cache-heavy — long agentic sessions, persistent context, big document bases read repeatedly — your real costs drop, by a reported 25 to 45% depending on the shape of the work. If your workload is novel, fresh reasoning where most tokens are new output rather than cached input, the cache cut barely touches you and you simply pay the ~20% verbosity premium. Same model, opposite cost outcomes, decided entirely by your token mix.

Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The knob that actually matters

The number nobody puts in the headline is the one a deployer should care about most: effort level. Fable 5.1 exposes five effort settings that span an 11-fold range in token usage, from about 13 million tokens at low effort to 144 million at max, scoring from 58 to 66 on the Index. Max effort — the 66 — is the most expensive corner of that range and rarely the rational default. Drop to "xhigh" and the model scores 65 at $2.72 per task, a dollar cheaper than max, and it still edges out Opus 5's 63 at a much smaller premium. Across the whole range Fable 5.1 sits on the intelligence-versus-tokens Pareto frontier, and its efficient floor is high — but the practical takeaway is that the leaderboard crown is set at the least economical setting, and most real deployments want to live a notch or two down, where you keep nearly all the intelligence and shed a real fraction of the cost. If you only remember one thing, remember that the effort dial, not the headline score, is where your budget is decided.

Amazon

AI token usage optimizer

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The asterisks that keep it honest

A few nuances that a careful reader shouldn't skip, and that I'd insist on for any vendor.

"Tops the leaderboard" is sometimes within the noise. On the agentic knowledge-work benchmarks, Fable 5.1's leads over Opus 5 are real on paper — 1,853 Elo to 1,824 on one, 1,694 to 1,685 on another — but Artificial Analysis says the first is within the confidence interval and the second is effectively a tie. The two models diverge in character rather than tier: Fable 5.1 is ahead on analytical quality and rubric correctness, and behind on presentation. So the crown is genuine at the top-line Index, but several of the underlying agentic margins are close enough that "meaningfully better than Opus 5 at real work" would be overstating what the data shows.

The accuracy record comes with more hallucination. Fable 5.1 attempts 93.4% of the questions on AA's knowledge benchmark, against Opus 5's 87.8%, and records the highest accuracy AA has measured at 67.2%. But that higher attempt rate cuts both ways: attempting more means getting more right and more wrong, and Fable 5.1 hallucinates more than its predecessor as a direct result. Whether that trade is good depends on whether your use case punishes a confident wrong answer more than it rewards a right one — which, for a lot of serious work, it does.

You're measuring the model plus its safety fallback. AA evaluated Fable 5.1 with Anthropic's default server-side routing, which sends safety-flagged requests to Opus 4.8 or Opus 5 instead; that fallback served about 4% of output tokens across the Index. It's a small share and doesn't change the picture much, but the thing being scored is Fable-5.1-with-fallback, not the base model in isolation — worth knowing if you're reasoning precisely about what produced the number.

Amazon

AI model verbosity control software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where I land

This is the same story I told about the frontier when Grok held its price flat, now pointed the other way. Intelligence and cost-per-task are separate axes, and the real competition is on the Pareto frontier between them, not at the single "smartest" point. Fable 5.1 pushes the intelligence ceiling to a genuine new high and cuts cache costs to defend the agentic-spend axis where the money actually is — and honestly, the 75% cache-read cut may matter more to most buyers than the three-point score, because it's a cost-to-serve move aimed at the workloads people actually run all day. The leaderboard crown is the marketing; the cache math and the effort dial are the substance.

So: yes, Fable 5.1 is the top of the Artificial Analysis Index, by a credible outside measure, with real broad gains — and it's more verbose and more expensive per task than its predecessor, leads Opus 5 by margins that are sometimes within noise, and buys its record accuracy partly with more hallucination. Both columns are true, and a grown-up reads both. If you're deploying it, ignore the number 66 and do the only calculation that matters: pick the effort level that holds the intelligence you need, and check whether your workload is cache-heavy enough for the price cut to catch you. Get those two right and this is a strong, sensibly-priced frontier model. Chase the headline setting and you'll pay top dollar for a leaderboard position you didn't need.


Analysis and opinion from a builder, founder, and post-labor economist running a local-first inference operation. All figures are Artificial Analysis's measurements (Intelligence Index v4.1.1), verified at time of writing against AA's published evaluation and reporting from 24/7 Wall St., officechai, cryptobriefing, and The Decoder: Index 66 (max) / 65 (xhigh); HLE 59.1%; Terminal-Bench v2.1 91.4%; SciCode 62.0%; ~$3.76 per task (max), ~20% above Fable 5 and ~1.6x Opus 5, driven by ~1.7x output tokens; cache read cut $1→$0.25 per 1M (75%), standard $10/$50 input/output unchanged; ~4% of output tokens served by Opus 4.8/5 fallback. Artificial Analysis disclosed it supported Anthropic with pre-release evaluation of the model; benchmark figures are third-party but should be read with that relationship in mind. This is analysis, not investment advice. Point-in-time as of 29 August 2026.

You May Also Like

Stargate Wisconsin: America’s Next AI Compute Frontier

AIThis post was created with the assistance of artificial intelligence (AI).By Thorsten…

Signal: Peak 2026 — Microsoft’s Anti-Mythos Weapon Includes Anthropic’s Own Models

AIThis post was created with the assistance of artificial intelligence (AI).Per an…

Redlines in Seconds, Citations in One Click: What Claude for Word Actually Does Inside a Contract

AIThis post was created with the assistance of artificial intelligence (AI).What can…

Twitch and e.l.f. Cosmetics Introduce Shoppable Livestream Ads

AIThis post was created with the assistance of artificial intelligence (AI).Introduction Livestreaming…