AIThis post was created with the assistance of artificial intelligence (AI).

The interesting story isn’t a smarter model. It’s the same intelligence at half the price, and what that does to the tasks you can afford to automate.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Two weeks after GPT‑6 Astra, OpenAI has filled out the family. GPT‑6 Sol and GPT‑6 Luna arrived on September 22, 2026, and the pitch isn’t a new capability ceiling. It’s the floor. Both models are priced at half their GPT‑5.6 predecessors, and the company frames them as models that help distribute the benefits of Astra’s intelligence by advancing the frontier on cost efficiency.

For anyone budgeting AI into a product or an operations workflow, that’s the more consequential kind of release. A smarter top-end model changes what’s possible. A cheaper mid-tier model changes what’s viable.

The prices

ModelInput (per 1M tokens)Output (per 1M tokens)
GPT‑6 Sol$2.00 (was $4)$10.00 (was $20)
GPT‑6 Luna$0.10 (was $0.20)$0.50 (was $1.20)

OpenAI says improvements in caching and inference let it serve these models at lower cost, with the savings passed on as a 50% price reduction against GPT‑5.6 promotional pricing. Cached input reads get a 90% discount, and the same 25% cache-write premium applies as before.

GPT‑6 Sol and Luna: half the price, about the same intelligence

OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.

GPT‑6 Sol
$4 / $20 → $2 / $10
GPT‑6 Luna
$0.20 / $1.20 → $0.10 / $0.50

Per 1M input / output tokens. Cached input reads keep the 90% discount.

Cost per task, halved

Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.

GPT‑5.6 Sol
$1.99
GPT‑6 Sol
$1.06
GPT‑5.6 Luna
$0.18
GPT‑6 Luna
$0.07

The effort dial moves cost more than the model choice

Model and effortIntelligence IndexCost per task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non‑reasoning)18$0.01

Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.

What got better, and what got worse

Better

  • Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
  • Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
  • OpenAI reports about half as many factual mistakes for Sol as its predecessor
  • Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing

Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.

Worse

  • GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
  • AA‑Briefcase v1.1: Luna down ~45 Elo
  • Coding Agent Index: Luna 41, down 2 points
  • Both models write more output tokens per task than their predecessors

Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.

What to do about it

Already on GPT‑5.6 Sol or Luna? The move is mostly a price cut. Re‑test first if your output is a document someone reads, not data a system consumes.
Shelved an automation on cost? Token prices halved and the effort dial adds another order of magnitude. Re‑run the business case.
Choosing between labs? The question is no longer which model is smartest, but which clears your quality bar at the lowest cost per task.
ThorstenMeyerAI.comSources: OpenAI (pricing, vendor benchmarks) and Artificial Analysis (independent evaluation and model pages). Figures as of 23 September 2026.

Astra remains the top of the range for work where you want the best result regardless of cost.

What the independent numbers say

Artificial Analysis published its evaluation the same day, and its summary is refreshingly blunt: cost per task halves while composite scores stay roughly level, with progress in some evaluations and regressions in others.

The cost side is unambiguous. Running its Intelligence Index costs $1.06 per task with GPT‑6 Sol at max effort, about 50% less than GPT‑5.6 Sol at $1.99, and $0.07 per task with Luna at max, roughly 60% less than its predecessor. That’s despite both models using slightly more output tokens per task, so the saving comes from price rather than efficiency.

On intelligence, Sol at max scores 48 on the Artificial Analysis Intelligence Index, well above the median of 25 for comparable models, with a 872k-token context window and 115 tokens per second of output. Luna at max scores 37 against a median of 12 in its price class, at 154 tokens per second with a 1M-token context window. For scale, Anthropic’s Claude Opus 5.5 took the top of the same index at 58 on the same day, with a 20% price cut of its own.

Coding is a split decision. In OpenAI’s Codex harness, Sol at max scores 57 on the Coding Agent Index, two points above its predecessor, with gains on Terminal-Bench 4.0 and SWE-Atlas-QnA, and it sits on the Pareto frontier of score versus cost. Luna at max scores 41, two points down, with lower scores on SWE-Atlas-QnA and DeepSWE v1.1.

Fewer confident wrong answers

The clearest quality improvement is in hallucination. On AA‑Omniscience, Sol at max cut its hallucination rate from 92% to 60%, and Luna from 93% to 77%.

The mechanism matters, though, and it’s a trade rather than a free win. Sol gets there partly by declining to answer more often: it attempts 83% of questions versus 99% for its predecessor, which cuts wrong answers by about a quarter while also lowering accuracy from 59% to 54%. On the composite index, which rewards correct answers, penalizes hallucinations and doesn’t punish refusals, Sol improves from 22 to 27 and Luna from ‑10 to 1.

OpenAI reports the same direction internally, saying Sol makes about half as many mistakes as its predecessor on a factuality evaluation built from real conversations where users had flagged errors.

If your use case is customer-facing answers or research support, a model that says “I don’t know” more often is usually the better colleague. If your use case is bulk extraction with a fixed expected output, more refusals means more empty cells.

The regression worth knowing about

Not everything improved. Artificial Analysis found regressions in two knowledge-work evaluations: on GDPval‑AA v2.1, adapted from OpenAI’s dataset of economically valuable tasks across 44 occupations, Sol dropped about 100 Elo points and Luna about 75, and Luna also fell around 45 Elo points on the AA‑Briefcase multi-week knowledge work benchmark.

The diagnosis is specific and, for business users, the most useful line in the whole analysis. After manually inspecting hundreds of outputs, the team attributes the regressions to reduced presentation quality and deliverables that omit rubric elements.

Read that next to OpenAI’s own release notes, which promise fewer low-value details and slightly shorter answers overall without losing substance. The same tuning that makes a model pleasanter in a coding chat appears to make it a weaker producer of complete, well-presented deliverables. If your workflow is “produce a finished document a client will read,” test before you switch. If it’s “make a decision inside a pipeline,” this probably doesn’t affect you.

The effort dial is the real cost lever

Both models expose reasoning effort levels, and the spread between them is larger than the spread between the two models. From the Artificial Analysis model pages:

Model and effortIntelligence IndexCost per index task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non-reasoning)18$0.01

Sol at low effort keeps about 70% of its max-effort score at roughly an eighth of the cost, because it produces 8.9M output tokens across the index instead of 77M. Luna at low effort costs less than half a cent per task.

Most teams pick a model and leave effort at the default. The table says the dial deserves as much attention as the model choice, and the right answer usually differs per task inside the same product.

Caching: the quiet cost story

Alongside token prices, OpenAI improved prompt caching so agents reuse more context, with 90% discounts on cached input reads, plus a caching dashboard and a diagnostics tool. Two changes matter for agent builders in particular: raising or lowering reasoning effort mid-conversation and turning tools on or off now both preserve earlier context for cache reuse, and explicit breakpoints let you choose where cached prefixes end.

The proof point offered is GitHub’s, which reports that these improvements cut the share of prompt tokens needing fresh processing by more than half across billions of requests.

For long-running agents, cache hit rate is often a bigger line item than the headline token price. Halving reprocessed input on a chatty agent can beat a 50% price cut.

Where OpenAI claims the wins

The vendor benchmarks are worth reading as cost-per-result claims rather than capability claims. On AutomationBench, a test of business workflows across 47 tools, Sol at xhigh effort scores 33.2% at $0.27 per task, above Claude Opus 5 at max effort on 26.9%, which cost 11.1 times as much per task. On DeepSWE v1.1, Sol at max scores 68.8%, within 1.1 points of Claude Fable 5’s best result at roughly 80% lower cost per task. On OSWorld 2.0, Sol at xhigh matches Claude Opus 5 at medium effort, 60.5% versus 60.3%, at about 80% lower cost.

Note the pattern: the comparisons are consistently “similar score, much lower cost,” not “higher score.” That’s the honest shape of this release, and it lines up with what the independent index found.

Availability

Sol and Luna are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, with Free and Go users getting Luna in the desktop app. They’re not in Chat yet. In the API they’re gpt-6-sol and gpt-6-luna.

What to do with this

If you’re running GPT‑5.6 Sol or Luna in production, the migration is mostly about price. Same-family, half the cost. Re-run your own evals first, especially if your output is a document someone reads rather than data another system consumes.

If you rejected an automation on cost grounds in the last year, the arithmetic has changed twice: token prices halved, and the effort dial gives you another order of magnitude at the low end. A workflow that didn’t pencil out at GPT‑5.6 prices may now.

If you’re choosing between labs, this release didn’t move the intelligence ceiling. Opus 5.5 leads the composite index while Sol leads on cost-efficiency in its price band. The right question is no longer “which model is smartest” but “which model clears my quality bar at the lowest cost per task,” and that’s a question only your own evaluation can answer.

The broader shift is worth sitting with. Two years ago, model launches were about what AI could newly do. This one is about what AI newly costs. That’s the sign of a technology moving from demonstration into infrastructure, and it’s the phase where the companies that measure carefully pull ahead of the ones that just upgrade.


Sources: OpenAI’s launch post, Artificial Analysis’s independent evaluation, and its individual model pages for Sol and Luna at each effort level. Figures as of September 23, 2026.

Amazon

AI language model API subscription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

cost-effective AI automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI token cost reduction software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI inference caching solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Compounding Error Problem — Why 99.9% Alignment Decays to 60% in 500 Generations

AIThis post was created with the assistance of artificial intelligence (AI).By Thorsten…

The unbundling of the budget app. Why a conversational finance surface absorbs what the personal-finance apps charge for, and what survives the absorption.

AIThis post was created with the assistance of artificial intelligence (AI).When Intuit…

AI Unveiled: A Data-Driven Briefing

AIThis post was created with the assistance of artificial intelligence (AI).This briefing…

Inside GPT-5.5: What OpenAI’s New System Card Actually Says About Its Frontier Model

AIThis post was created with the assistance of artificial intelligence (AI).Published April…