OpenAI’s new models make capable AI substantially cheaper. The harder question is how much of that saving survives once the work has to be checked, corrected and accepted.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
By Thorsten Meyer | 23 September 2026 | Reality Check
An AI task is finished when somebody can use the result.
That sounds obvious. Yet much of the discussion around model releases still stops at the price of generating it. A cheaper answer becomes a productivity gain before anyone has checked whether the spreadsheet reconciles, the software change passes review or the report contains everything the brief required.
GPT-6 Sol and Luna make that distinction economically important.
OpenAI has cut Sol’s API prices to $2 per million input tokens and $10 per million output tokens. Luna falls to $0.10 and $0.50 respectively. Against the promotional GPT-5.6 prices in the announcement, Sol’s rates are halved; Luna’s input rate is halved and its output rate falls by approximately 58%. [1]
Those are meaningful reductions. They expand the range of work worth attempting with AI.
ThorstenMeyerAI.com / Reality Check
GPT-6 Sol & Luna
Cheaper intelligence. The review bill remains.
01 The price of generation
SOL
Standard API token rates · USD
LUNA
Standard API token rates · USD
02 Reasoning effort changes the bill
Artificial Analysis Intelligence Index: score and weighted cost per task.
| Effort | LUNA | SOL | ||
|---|---|---|---|---|
| Score | Cost / task | Score | Cost / task | |
| Non-reasoning | 18 | $0.01 | 28 | $0.33 |
| Low | 21 | $0.0045 | 34 | $0.13 |
| Medium | 29 | $0.02 | 40 | $0.25 |
| High | 32 | $0.03 | 43 | $0.37 |
| xhigh | 34 | $0.04 | 44 | $0.53 |
| Max | 37 | $0.07 | 48 | $1.06 |
Same index score ≠ the same abilities on your workload.
03 A 50% API saving can disappear in review
Illustrative scenario · reviewer cost: $45/hour · not measured model results.
Baseline
Model: $1.00
Review: 4 min = $3.00
$4.00 TotalCheaper model
Model: $0.50
Review: 4 min = $3.00
$3.50 TotalOne extra minute
Model: $0.50
Review: 5 min = $3.75
$4.25 TotalMeasure cost per accepted result
Model + tools + review + rework spendingdivided by the number of accepted results
Sources: OpenAI launch announcement and Artificial Analysis.
All twelve model sources
Snapshot: 23 September 2026. Benchmark task costs are not business-workflow forecasts. Scores are not success rates. Values retain the source’s displayed precision.
But the practical question is what happens after the attempt. My reading of this release is that businesses should start measuring the cost of an accepted result with the same discipline they apply to the cost of generating one.
AI productivity tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the independent assessment actually says
Artificial Analysis describes broadly level aggregate intelligence, with uneven changes underneath. In its Coding Agent Index, Sol at maximum effort rises two points to 57; Luna falls two points to 41. Sol’s coding evaluation costs $2.99 per task, roughly half its predecessor’s cost. [2]
The more consequential warning concerns knowledge work. Both models regress on GDPval-AA v2.1. Luna also declines on AA-Briefcase v1.1, while Sol stays level. Artificial Analysis attributes the regressions it inspected to weaker presentation and missing rubric requirements. [2]
Its factuality results also need careful reading. Sol’s AA-Omniscience hallucination rate falls from 92% to 60%, partly because it attempts fewer questions; accuracy falls from 59% to 54%. Luna’s hallucination rate falls from 93% to 77%, with accuracy broadly unchanged. These are benchmark-specific measurements, not everyday error probabilities. [2]
That is enough evidence to reject an automatic upgrade policy. An organization should retain examples of work its existing setup does well and run the replacement against the same acceptance criteria. A cheaper model earns its place by delivering usable work under those conditions.
As an affiliate, we earn on qualifying purchases.
There are twelve configurations to compare
The model name is only the first purchasing decision. Reasoning effort changes the economics substantially.
Artificial Analysis lists six configurations for each model. Its unsuffixed Sol and Luna pages represent max effort. They should not be confused with non-reasoning or an unspecified default.
The following snapshot uses the Intelligence Index scores and weighted cost per Intelligence Index task displayed on 23 September 2026. Dollar values retain the source’s displayed precision. These are benchmark costs, not forecasts for a business workflow. [3–14]
Reasoning configuration Luna: Intelligence Index Luna: cost per task Sol: Intelligence Index Sol: cost per task Non-reasoning 18 $0.01 28 $0.33 Low 21 $0.0045 34 $0.13 Medium 29 $0.02 40 $0.25 High 32 $0.03 43 $0.37 xhigh 34 $0.04 44 $0.53 Max 37 $0.07 48 $1.06
Three comparisons deserve attention.

First, Luna at xhigh and Sol at low both score 34. Their displayed costs are $0.04 and $0.13. On this aggregate measure, allocating more reasoning to the cheaper model produces the same score at about 69% lower cost. Equal index scores do not establish equal performance on the tasks a particular business needs. They do establish a useful comparison to test. [7, 10]
Second, Sol’s move from high to xhigh raises the score from 43 to 44 while the displayed cost rises from $0.37 to $0.53 — approximately 43%. That extra point may be valuable on a difficult workload. The table provides no reason to buy it automatically for every request. [12, 13]
Third, non-reasoning is not the cheapest configuration in either family on these measurements. Low effort has a higher score and a lower weighted task cost. That result concerns this evaluation mix; it does not guarantee that adding reasoning reduces the bill for a simple classification request. [3, 4, 9, 10]
The useful procurement question becomes: which model and effort setting meets the acceptance threshold for this kind of work?
As an affiliate, we earn on qualifying purchases.
Token prices are a starting point
At the listed rates, Luna’s input and output tokens each cost one-twentieth as much as Sol’s. An illustrative request using 10,000 uncached input tokens and 2,000 billable output tokens would therefore cost $0.002 on Luna and $0.04 on Sol. This calculation excludes tools, retries and any additional billable tokens. It assumes identical usage. [4, 10]
Actual usage can differ. The cost of a workflow depends on how often the model calls tools, how much context it processes, how long it reasons and whether somebody sends it back to try again.
Artificial Analysis’s task-cost methodology accounts for input, cache reads, cache writes, reasoning and answer tokens, with costs weighted across the evaluations. It is more informative than comparing a single token rate, while still describing a particular benchmark workload. [4]
Caching adds another variable. OpenAI advertises a 90% discount on cached input reads and says changing reasoning effort or tool availability can now preserve earlier context for reuse. That can matter for recurring agent workflows. It is not a 90% reduction in the total bill. [1]
An organization that repeatedly supplies the same background material should measure its actual cache usage. An organization whose work consists largely of new documents should not budget as if every request reuses a previous one.
As an affiliate, we earn on qualifying purchases.
The saving has to survive human review
Consider a deliberately simplified example. These figures illustrate the economics; they are not measured Sol or Luna performance.
Suppose an AI workflow costs $1 per attempt and requires four minutes of review. At an assumed fully loaded reviewer cost of $45 an hour, that review costs $3. The combined cost is $4.
Halve the model bill and leave review unchanged: the total falls to $3.50. A 50% inference saving produces a 12.5% saving across the two cost components.
Now suppose the cheaper setup needs one additional minute of review. That minute costs $0.75. The total rises to $4.25, exceeding the original cost despite the smaller API bill.
The point is not that lower-priced models require more checking. That has to be measured. The point is that even a small change in review time can dominate a large percentage change in inference spending.
For an operational pilot, I would track one principal measure:
Cost per accepted result = total model, tool, review and rework spending ÷ accepted results.
Keep completion time and the severity of mistakes alongside it. A workflow can be cheap and still fail because the answer arrives too late, or because the remaining mistakes are unacceptable.
This also gives businesses a sensible way to compare human-assisted and automated workflows. Count the work required to reach the same standard in both cases.
Route work according to evidence
Sol and Luna invite a practical experiment: divide a workload into categories, then test whether different configurations deserve different roles.
For constrained extraction, tagging or routine transformations, begin by testing Luna at low or medium effort. Use tasks where correctness can be checked against source material or explicit rules.
For work requiring more judgment, compare Luna at higher effort with Sol at low or medium. The aggregate results make this an interesting boundary to examine, but they cannot settle it for you.
For complex work, test Sol at high, xhigh and max against the same brief. Record which additional requirements are satisfied when the reasoning budget increases. A higher setting should justify itself through better outcomes or less rework.
These are proposed starting points for an evaluation, not claims that either model has been validated for every task in those categories.
Escalation also needs a rule. Missing evidence, failed tests and an unmet requirement are useful signals. A model’s confident tone is a poor acceptance test.
And a second model should not automatically be treated as an independent auditor. Where possible, verification should reach back to the source document, the calculation, the software test or a qualified reviewer. Repeated agreement is less valuable than a check that can actually detect the error.
Cheaper AI changes who can experiment
There is a broader economic implication here.
When the cost of attempting a task falls, a smaller organization can afford to test more uses. Work that never justified custom automation may become worth examining: preparing a first draft of a recurring report, structuring a neglected archive or checking whether incoming records contain required fields.
That creates an opportunity for businesses whose advantage lies in knowing the work closely. They can define what a usable result looks like, identify the exceptions and see where review time disappears.
It also creates a management temptation: treat the lower generation price as evidence that the entire process is ready to scale. An inexpensive system can generate a considerable backlog of unfinished work.
The next constraint may be the organization’s capacity to verify and act on what it produces. That is an inference about deployment economics, not a finding established by the model benchmarks.
For the debate about AI and employment, this distinction matters. These sources do not establish a number of jobs that can be eliminated. They supply evidence about model performance and cost. Translating that into labor demand requires knowing which tasks become usable, what supervision remains and whether the lower cost creates additional demand.
The release deserves a serious trial
GPT-6 Sol and Luna give organizations a reason to revisit their model budgets and their routing decisions. The twelve configurations show how much is hidden by a single product name.

My conclusion is that the strongest adoption case will come from a measured workflow: a defined brief, a visible acceptance threshold, a record of corrections and an honest accounting of the time people spend finishing the work.
The price cut makes more attempts affordable.
The business case begins when more of those attempts become work someone can use.
Sources and methodology
Sources checked on 23 September 2026. Model-page values are a dated snapshot and may change. The Artificial Analysis model pages identify Intelligence Index v4.3.2; its launch article refers to v4.3. Launch comparisons and current model-page figures are identified separately above. Index scores are aggregate scores, not percentages of workplace tasks completed successfully. Calculations and deployment proposals are this article’s analysis; no independent hands-on testing is claimed.
- OpenAI — Introducing GPT-6 Sol and Luna: API pricing and caching claims.
- Artificial Analysis — GPT-6 Sol and Luna push the cost efficiency frontier: launch assessment, predecessor comparisons, coding, knowledge-work and factuality results.
- GPT-6 Luna — non-reasoning
- GPT-6 Luna — low
- GPT-6 Luna — medium
- GPT-6 Luna — high
- GPT-6 Luna — xhigh
- GPT-6 Luna — max
- GPT-6 Sol — non-reasoning
- GPT-6 Sol — low
- GPT-6 Sol — medium
- GPT-6 Sol — high
- GPT-6 Sol — xhigh
- GPT-6 Sol — max
This article was created with the assistance of artificial intelligence.
Publishing details
Suggested category: Reality Check
Suggested slug: gpt-6-sol-luna-cost-of-accepted-work
SEO title: GPT-6 Sol and Luna: Cheaper AI, but What Does Work Cost?
Meta description: GPT-6 Sol and Luna cut AI costs. Twelve configurations reveal the trade-offs—and why review time determines the real cost of accepted work.
Excerpt: Sol and Luna make capable AI cheaper to use. The independent results are mixed, and the business case depends on what survives review. A look at all twelve configurations and the economics of accepted work.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
