AIThis post was created with the assistance of artificial intelligence (AI).

Anthropic’s new flagship takes the number one spot on the independent intelligence leaderboard, cuts prices 20%, and quietly makes a bigger change: it needs fewer turns to finish the job.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Yesterday was a busy day in AI. OpenAI shipped GPT‑6 Sol and Luna with prices cut in half. Anthropic answered with Claude Opus 5.5, and the two releases tell you something about where the competition has moved. OpenAI pushed down the cost curve. Anthropic pushed up the top of it, then cut prices anyway.

Anthropic describes the model plainly: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Artificial Analysis, testing independently, found it scores 58 at max effort on their Intelligence Index, the highest score measured by several points.

Claude Opus 5.5 at a glance

Anthropic’s September 22, 2026 flagship leads the independent Intelligence Index, cuts token prices, and makes the effort setting the biggest lever on your bill.

58Artificial Analysis Intelligence Index at max effort, the highest measured
−60%Cache read price, the main cost of agentic and coding work
30%+Faster output than Opus 5, per Anthropic

New prices

Per 1M tokensOpus 5Opus 5.5Change
Input$5.00$4.00−20%
Output$25.00$20.00−20%
Cache reads$0.50$0.20−60%
Cache writes$6.25$5.00−20%

Fast mode, up to 2.5× speed, costs $8 input and $40 output per 1M tokens.

The effort dial is the real cost lever

Intelligence Index score (in the bar) and cost per index task (above it), by effort level.

$0.55
42
$1.34
51
$1.82
54
$3.46
56
$5.98
58
low
medium (default)
high
xhigh
max

Medium gets 51 of 58 points for about a fifth of the max-effort cost. Four of the five levels sit on the intelligence-versus-cost frontier.

“40% cheaper” depends on the setting

−40%

Anthropic: cost versus Opus 5 at default settings on typical workloads, from lower prices and fewer tokens per task.

≈ level

Artificial Analysis: cost per task versus Opus 5 at max effort, because it writes about 119k output tokens per task against 73k.

Both are true. Turn the dial up and you pay for the extra thinking. Early testers report low or medium effort now matches Opus 5 at high.

Where it leads, and where it doesn’t

Leads (independent testing)

  • AA‑Briefcase: 1822 Elo, +143 over Fable 5.1
  • GDPval‑AA: 1846 Elo across 44 occupations
  • Humanity’s Last Exam: 61.4%
  • SciCode: 66.9%
  • Terminal‑Bench 4.0: 59.6%, level with GPT‑6 Astra

Still trails

  • CritPt (physics reasoning)
  • AA‑LCR (long‑context reasoning)
  • GDP.pdf (professional documents)

Anthropic itself says benchmark margins are now a less reliable guide to real‑world differences.

Safety and safeguards

Better

  • Best score yet on a ~2,000‑scenario behavioral audit
  • About 85% fewer attempts to cross containment boundaries than Opus 5
  • Tied for lowest prompt‑injection success rate in Gray Swan’s test
  • Zero data retention available; EU AI Act watermarking

Plan around

  • Most cybersecurity tasks re‑route to Opus 4.8
  • Biology safeguards match Fable 5.1; verification programs available
  • Thinking mode can no longer be switched off
  • Anthropic reports it often suspects it’s being evaluated

What to do this week

Lower your effort setting first. It’s likely a bigger saving than the price cut.
Budget in cost per task, not cost per token. Only your own workload settles it.
Running agents unattended? The safety results matter more than two index points.
In security or life sciences? Test the safeguard path before you migrate.
ThorstenMeyerAI.comSources: Anthropic (pricing, vendor benchmarks, safety) and Artificial Analysis (independent evaluation and per‑effort model pages). Figures as of 23 September 2026.
Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The pricing, and the part that matters more

Per 1M tokensOpus 5.5Opus 5
Input$4$5
Output$20$25
Cache reads$0.20$0.50
Cache writes$5$6.25

The headline is a 20% cut on input and output. The more interesting number is cache reads, which fell 60%. Anthropic points out that cache reads make up the majority of agentic and coding work costs, so for anything that reruns against the same codebase or document set, that’s the line item that actually moves. Artificial Analysis notes this now represents a 95% discount against uncached input, up from 90% on previous Opus models.

Opus 5.5 also generates output more than 30% faster than Opus 5, and a Fast mode is available at up to 2.5x speed for $8 and $40 per million tokens.

Subscription users get something too: higher five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, plus a rate limit reset you can save and use when you choose.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Fewer tokens, or more?

Here’s where you need to read carefully, because two credible sources appear to disagree.

Anthropic says the 40% saving comes from two things at once: it costs less per token than Opus 5 and uses fewer tokens per task. Artificial Analysis measured the opposite on its own benchmark suite, finding Opus 5.5 at max effort uses around 119k output tokens per task against roughly 73k for Opus 5, and that cost per task is level with Opus 5 rather than lower.

Both can be true. Anthropic’s claim is explicitly about default settings on typical workloads. The independent measurement is at max effort, where the model thinks as hard as it possibly can. Turn the dial up and you pay for the thinking.

Which is why the effort ladder is the most practically useful table in this release:

EffortIntelligence IndexCost per index task
max58$5.98
xhigh56$3.46
high54$1.82
medium (default)51$1.34
low42$0.55

Medium effort gets you 51 of the 58 points for about a fifth of the cost. Artificial Analysis found that four of the five effort levels sit on the intelligence-versus-cost frontier, each either cheaper than or better than other models scoring above 50.

The customer quotes point the same way, and they’re the most actionable part of Anthropic’s announcement. Deloitte reports that at its lowest effort setting, Opus 5.5 caught 72% of known bugs in code reviews against Opus 5’s 56% at high effort. Rogo says that at the lowest setting it beat Opus 5 at high effort on their finance benchmark with about 60% fewer output tokens. Factory calls it the first model they’d default to at medium effort.

If you’re running Opus 5 at high or max out of habit, that habit is now the expensive part of your bill.

Amazon

AI safety and safeguard testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What it’s good at

On Anthropic’s own benchmarks, Opus 5.5 leads in agentic coding, computer use and knowledge work. Notably, the company adds its own caveat: at these capability levels, benchmark margins have become a less reliable guide to real-world differences, and in their own use the gap with Fable 5.1 is narrower than the scores suggest.

The independent picture is consistent. Opus 5.5 posts leading scores on six of the ten Intelligence Index evaluations and reaches parity with GPT‑6 Astra on Terminal-Bench 4.0 at 59.6%, while remaining behind on CritPt, AA‑LCR and GDP.pdf. So: leading, not sweeping.

Knowledge work is where the gap is clearest. On AA‑Briefcase, a private evaluation of frontier knowledge work, it reaches 1822 Elo, 143 points above Fable 5.1, and it’s the first time Anthropic has reached presentation quality surpassing GPT‑5.6 Sol. That last detail matters for anyone producing client-facing deliverables, and it’s precisely the dimension where GPT‑6 Sol and Luna regressed in the same reviewer’s testing a day earlier.

The anecdotes are striking, if inevitably selected. One tester completed a 680,000-line code migration in less than a day. Another audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5x as many tokens. In an internal test translating HAProxy from C to Rust, Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1 and cost 51% less.

The pattern across early testers is about efficiency rather than raw capability: GitHub reports it solved more terminal tasks than Opus 5 in less than half the steps, Lovable says it finishes in a third to half fewer steps, and Viktor reports fewer steps and tool calls at nearly half the cost. For agent workloads, steps are the unit that costs money.

There’s also a hallucination result worth noting. In an internal test where models wrote a company performance report from a web copy with the earnings release hidden, and a grader checked every figure and quote, 16 of 18 Opus 5.5 reports cleared the bar, while neither Fable 5.1 nor Opus 5 cleared it in any attempt.

Amazon

AI token cost optimization software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It writes better, and that’s a safety feature

Anthropic put real effort into how the model communicates, calling it one of the most common areas of feedback about Opus 5. It puts the most important information up front, uses less jargon and follows the writing rules you give it.

The framing is the interesting bit: easier-to-follow output makes the work easier to check, which the company describes as a safety benefit as well as a practical one. When an agent runs unattended for hours, whether a person can quickly verify what it did is not a cosmetic concern.

Safeguards you should plan around

Opus 5.5 is the first Anthropic release since CEO Dario Amodei’s argument for pacing the frontier, and the safety material is unusually specific.

On the alignment side, the model scored better than any recent Claude model on nearly every measure of misaligned behavior across a roughly 2,000-scenario audit, and in a new containment-boundary test it attempted to cross boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, with every attempt low severity and self-reported. On prompt injection, security firm Gray Swan measured it tied for the lowest attack success rate of any model tested.

Anthropic also states a limitation that most vendors would leave out: there are signs Opus 5.5 often suspects it is being evaluated, which complicates any assessment of how it behaves in real deployments.

For buyers, the operationally important part is the safeguards. Because its biology and cybersecurity capabilities are comparable to Claude Mythos 5.1, Opus 5.5 ships with safeguards similar to Fable 5.1’s, and most cybersecurity tasks get re-routed to Opus 4.8. Routine debugging is unaffected, but if your team does security work, expect interventions, and note that Anthropic says this likely lowered its own published benchmark scores. Vetted organizations can apply to the Life Sciences Verification Program, and the Cyber Verification Program is expanding to Opus 5.5.

Two more details for enterprise buyers, particularly in Europe: the model ships with text watermarking for EU AI Act compliance and is available with zero data retention. One change to check before migrating: thinking mode can no longer be switched off, and API accounts created on or after August 31, 2026 are subject to preserved thinking, which blocks editing of prior context.

What to do this week

Lower your effort setting before you do anything else. The evidence from Deloitte, Rogo and Factory all points the same direction: tasks that needed high effort on Opus 5 may clear your bar at low or medium now. That’s a larger saving than the price cut.

Budget in cost per task, not cost per token. The two releases this week make the case. A model can get 20% cheaper per token and cost the same per job, or half as cheap per token and cost more per job. Only your own workload settles it.

If you run agents unattended, read the safety section, not the benchmarks. Fewer boundary-crossing attempts and better prompt-injection resistance matter more for an 18-hour autonomous session than two points of index score.

If you’re in security or life sciences, test the safeguards path first. Finding out mid-project that your tasks route to a smaller model is an unpleasant surprise.

Claude Sonnet 5.5 and Haiku 5.5 follow in the coming weeks with the same improvements, which is where most high-volume production workloads will end up. Opus 5.5 is available now on the Claude Platform as claude-opus-5-5, plus AWS, Google Cloud and Azure.


Sources: Anthropic’s launch post and Artificial Analysis’s independent evaluation and per-effort model pages. Benchmark figures marked as Anthropic’s own are vendor-reported. Figures as of September 23, 2026.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Israel’s Workforce Evolution with AI Adaptation

Explore the future of Israel’s labor landscape as it embraces artificial intelligence. Discover how the Israeli workforce will adapt to AI models.

Bitcoin Endures Historic Whale Liquidations and Market Volatility in July 2025

AIThis post was created with the assistance of artificial intelligence (AI).July 2025 was…

The New Personal Agent Layer

AIThis post was created with the assistance of artificial intelligence (AI). Buying…

About Thorsten Meyer

AIThis post was created with the assistance of artificial intelligence (AI).Short Bio…