AIThis post was created with the assistance of artificial intelligence (AI).

Anthropic’s latest model leads the Artificial Analysis Intelligence Index. The more useful business question is how much of that capability you need to buy on every task.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

By Thorsten Meyer | 23 September 2026 | Reality Check

Claude Opus 5.5 arrived on 22 September with a straightforward proposition from Anthropic: stronger performance and lower operating costs. Artificial Analysis supplies independent evidence for the capability claim, placing the model first on its Intelligence Index at maximum effort, with a score of 58. [1, 7]

That makes this a significant release. It also makes an expensive deployment mistake easier to justify.

A team sees a new leader, selects the highest reasoning setting and assumes it has bought the sensible default. But the five configurations tell a more complicated story. On Artificial Analysis’s current model pages, medium effort scores 51 at $1.34 per benchmark task. Max scores 58 at $5.98. [4, 7]

Seven additional index points cost roughly four and a half times as much.

Those points can be worth paying for. The question is which tasks deserve them—and whether the organization has a way to find out.

ThorstenMeyerAI.com / Reality Check

Claude Opus 5.5

The benchmark leader. Five different budgets.

01 What does maximum effort buy?

MEDIUM

51Intelligence
Index score

$1.34 per benchmark task

MAX

58Intelligence
Index score

$5.98 per benchmark task

4.46×
the cost of medium, for 7 additional index points

Calculated from displayed benchmark costs. Extra points are not a proportional measure of business value.

02 Compare all five settings

Adaptive reasoning · default fallback enabled in every configuration.

Artificial Analysis Intelligence Index v4.3.2 · USD · 23 September 2026. Swipe horizontally on narrow screens.
EffortIndex scoreCost / taskvs. medium
Low42$0.550.41×
Medium51$1.341.00×
High54$1.821.36×
xhigh56$3.462.58×
Max58$5.984.46×

Weighted cost per Intelligence Index task. Scores are not task success rates.

03 Read the claims at the right level

  • Token pricing: $4 input / $20 output per million tokens. Cache reads: $0.20 per million.
  • Anthropic’s cost claim: approximately 40% lower cost than Opus 5 on typical workloads at default settings.
  • Independent max-effort result: Artificial Analysis reports roughly level cost per task versus Opus 5, with more output tokens.
  • Different settings, different workloads: neither comparison guarantees your production savings.

A practical starting point

Test medium and high. Escalate where the extra effort pays.

Measure accepted results, correction time, retries and the complete workflow bill. This is an evaluation proposal, not a benchmark finding.

Sources: Anthropic launch announcement · Artificial Analysis launch assessment

Five model sources

Snapshot: 23 September 2026. All configurations include default fallback; results describe that evaluated setup. Benchmark task costs are not production quotes. Relative costs use rounded displayed values.

Thorsten Meyer AIBuy the effort your workflow needs
Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The strongest signal is professional work

Artificial Analysis reports leading results on six of the ten Intelligence Index evaluations. Its assessment highlights agentic knowledge work: Opus 5.5 reaches 1,822 Elo on AA-Briefcase, 143 ahead of Fable 5.1, leading in analytical quality and presentation. It remains slightly behind Fable on that evaluation’s rubric-based scoring. [2]

That combination matters. In professional work, the answer and the deliverable belong to the same job. A correct conclusion buried in a confusing report can still require substantial human effort before anyone can use it. A clear document that omits a required analysis creates a different kind of failure.

The useful interpretation is that Opus 5.5 deserves a serious trial on work where both the reasoning and its presentation matter. The separate rubric result is a reminder to inspect completeness as well as polish.

A business evaluating these models should therefore retain the brief alongside the finished output. Did the work answer every required question? Are the assumptions visible? Can another person follow the calculation? Can the recipient use the deliverable without rebuilding it?

These are measurable properties of work. They also provide a better basis for a purchasing decision than whether an answer feels impressive.

Amazon

AI cost analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five effort settings, five different bills

The model pages supplied for this article all specify adaptive reasoning with default fallback. Their displayed Intelligence Index v4.3.2 results are summarized below. The unsuffixed model page is max, not an unspecified default. [3–7]

Effort settingIntelligence IndexWeighted cost per index taskCost relative to medium
Low42$0.550.41×
Medium51$1.341.00×
High54$1.821.36×
xhigh56$3.462.58×
Max58$5.984.46×

Snapshot: 23 September 2026. Costs are USD and retain the source’s displayed precision. Relative costs are calculated from those rounded values. Index scores are not percentages of workplace tasks completed successfully.

Low to medium buys nine points for an additional $0.79 per benchmark task. Low is the least expensive of these configurations, but its 42-point score is a materially different result from the headline 58. [3, 4]

Medium to high adds three points for $0.48, a cost increase of approximately 36%. High therefore deserves to sit beside medium in an initial evaluation, particularly where mistakes create appreciable rework. [4, 5]

High to xhigh adds two points while increasing the displayed cost from $1.82 to $3.46—about 90%. The next step, xhigh to max, adds another two points for $2.52 more per task, approximately 73% extra. [5–7]

These comparisons do not establish a universal optimum. An aggregate score can conceal a large gain on the one capability a business needs. Equally, an organization may pay for gains on tasks it never performs.

The table supports a narrower conclusion: effort is a budget decision that should be evaluated separately from the choice of model.

My starting proposal would be to test medium and high on representative work, then reserve xhigh and max for categories where they demonstrably improve acceptance, reduce correction time or solve failures that the lower settings leave unresolved.

Amazon

professional AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The price cut is real; the workload saving varies

Anthropic cuts standard input and output token prices by 20%, and cache-read rates by 60%. Its claim of approximately 40% lower cost concerns typical workloads at default settings; the launch page identifies medium as the default effort level. [1]

A reduction in token prices and a reduction in the cost of finishing a task are different measurements.

Artificial Analysis reports that max-effort Opus 5.5 uses roughly 119,000 output tokens per Intelligence Index task, against about 73,000 for Opus 5. Its cost per task is approximately level with its predecessor despite that higher token use. [2]

These findings need not contradict Anthropic’s default-setting claim. They describe different workloads and reasoning settings. Treating either as a guarantee for every deployment would go beyond the evidence.

For budgeting, the first step is to record the configuration that produced the claimed saving. The second is to reproduce the comparison on the organization’s own work. A forecast should say which effort setting it assumes, how much context is reused and how often a task needs another attempt.

Without that information, a percentage saving is a description of someone else’s experiment.

Amazon

enterprise AI workflow evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Caching rewards repeated context

Opus 5.5’s $0.20 cache-read price is 95% below its $4 uncached-input rate. [5] Artificial Analysis lists five-minute cache writes at $5 per million tokens. [2]

For an illustrative calculation, one million eligible cached input tokens therefore cost $0.20 to read, versus $4 as uncached input. That is $3.80 less for that read. It does not include writing the cache, generating output or paying for tools. [5]

The business implication is conditional but useful. A workflow that repeatedly consults the same substantial background material has a different cost structure from one that mostly processes fresh material. An engineering agent revisiting a repository may have opportunities for reuse; a stream of unrelated requests may have fewer.

Those are workload examples, not measured Opus 5.5 savings. What matters is the actual proportion of requests that qualify for reuse and the complete bill over the workflow’s lifetime.

A procurement comparison based only on the output-token rate will miss that distinction. A comparison based only on the cache discount will miss the rest of the spending.

“With fallback” belongs in the headline evidence

There is an important qualification attached to every configuration in the table: default fallback is enabled. [3–7]

Anthropic explains that, in its benchmark evaluations, safeguard interventions routed cybersecurity tasks to Opus 4.8, and biology and frontier-model-development tasks to Opus 5. It separately notes that the cited AutomationBench runs used no fallback. [1]

This means readers should treat the reported results as results of the stated evaluated setup. They should not assume every task was completed solely by the same underlying model.

For deployment, that raises a practical question: does the product or API configuration being purchased behave like the configuration being evaluated?

A team should record routing behavior where it is exposed, examine what happens when a safeguard intervenes and check whether the fallback completes the original task satisfactorily. These observations belong beside quality and cost in a pilot.

Fallback can be a useful part of a service. It also changes what a model label tells you. The model name alone does not fully describe the system doing the work.

The premium can pay for itself surprisingly quickly

A higher inference bill is not automatically a worse business outcome.

Consider another illustrative calculation, using the benchmark cost difference only as an arithmetic example. Max costs $4.64 more per task than medium in the table. At an assumed fully loaded reviewer cost of $60 per hour, $4.64 buys four minutes and 38 seconds of human time.

If max reliably saved more review time than that on a comparable real task, the additional model spending could pay for itself through labor savings alone. If it saved no review time and changed nothing about acceptance, the extra spending would need another justification.

This is not a prediction that max saves those minutes. It is a way to express the threshold a pilot should measure. Actual inference costs must replace the benchmark averages before using the calculation operationally.

The same logic applies to elapsed time and retries. A lower-cost first attempt that repeatedly stalls may lose its advantage. An expensive attempt that produces a polished but incorrect result may be worse still.

For each task category, measure the total cost of reaching an accepted result, including model calls, tools, corrections and human review. Retain the failed attempts in the accounting. Otherwise, the workflow that generates the most waste can look artificially cheap.

Give the new leader a bounded job

Opus 5.5’s results justify attention. They do not remove the need to specify what a deployment must accomplish.

A useful first trial would take a small set of recurring tasks with known acceptance criteria and run medium and high against the current workflow. Include awkward cases: incomplete source material, conflicting instructions and briefs with several deliverables. Review the outputs without letting the model label decide the grade.

Then escalate the failures to higher effort. Record whether the additional spending fixes the failure, produces a different failure or merely extends the answer. This reveals where the top settings add value and where a better brief, a different tool or human judgment is the missing ingredient.

For work that can change external systems, keep permissions proportionate to the task being tested. A stronger benchmark result does not answer which actions a business wants to delegate.

The broader economic opportunity is an organization that can allocate expensive reasoning deliberately. Routine work gets a measured budget. Difficult work gets more when the evidence supports it. Reviewers spend their time on unresolved questions rather than supervising every step equally.

That is the practical promise I would test in this release.

Opus 5.5 has earned a place at the front of the evaluation queue. Maximum effort still has to earn its place in the workflow.


Sources and methodology

Sources checked on 23 September 2026. The launch date is 22 September 2026. Rankings and model-page values are a dated snapshot. All five configurations use default fallback. Artificial Analysis reports weighted costs for its evaluation mix; these are neither quotes for a business workflow nor costs per successfully accepted result. Index scores and Elo ratings are different measures. Calculations and deployment proposals are this article’s analysis. No independent hands-on testing is claimed.

  1. Anthropic — Introducing Claude Opus 5.5: launch date, pricing, default-setting cost claim, default effort and benchmark fallback notes.
  2. Artificial Analysis — Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index: launch assessment, knowledge-work results, token use and caching details.
  3. Claude Opus 5.5 — low, with default fallback
  4. Claude Opus 5.5 — medium, with default fallback
  5. Claude Opus 5.5 — high, with default fallback
  6. Claude Opus 5.5 — xhigh, with default fallback
  7. Claude Opus 5.5 — max, with default fallback

This article was created with the assistance of artificial intelligence.


Publishing details

Suggested category: Reality Check
Suggested slug: claude-opus-5-5-benchmarks-medium-vs-max
SEO title: Claude Opus 5.5: Benchmarks, Pricing and the Cost of Max
Meta description: Claude Opus 5.5 leads the Intelligence Index. Compare five effort settings, caching costs and fallback caveats before making max your default.
Excerpt: Opus 5.5 takes the benchmark lead, but max effort costs about 4.5 times medium in Artificial Analysis’s evaluation. The business case depends on where extra reasoning improves accepted work.
Featured image alt text: A copper-lit computational core above five rising architectural platforms, illustrating Claude Opus 5.5 and the choice of reasoning effort.
Infographic placement: After “Five effort settings, five different bills.” Paste the companion HTML file into a WordPress Custom HTML block.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ASI-ARCH: A New Era of Autonomous AI Research

AIThis post was created with the assistance of artificial intelligence (AI).The AI…

Lovable Was the Most Copied AI Product of 2025. Then a Lobster Changed Everything

AIThis post was created with the assistance of artificial intelligence (AI).A new…

Three Days at the Frontier: Washington Suspends Fable 5 and Mythos 5

AIThis post was created with the assistance of artificial intelligence (AI).Three days.…

Elon Musk’s xAI Pursues $12B in Private Funding

Discover how Elon Musk’s xAI is reshaping the future by seeking an impressive $12B in private credit funding for its groundbreaking initiatives.