This week OpenAI published 722 mathematical manuscripts. The average result took about three hours of compute to produce. The field’s response to a single earlier result from the same programme, a counterexample to an old Erdős conjecture, was a careful verification by five of the world’s leading mathematicians.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That ratio is the story of the next decade of work, and it isn’t confined to mathematics.
AI has collapsed the cost of producing a result. It hasn’t touched the cost of trusting one. In some fields it has made trust more expensive. Producing is now cheap and abundant. Checking is slow, human, and scarce.
For a post-labor economy, that asymmetry matters more than any benchmark. It tells you which work is disappearing, which work is quietly becoming the bottleneck, and which jobs nobody is training for.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Three fields, one pattern
Mathematics. OpenAI’s model was posed about 4,000 problems and produced 722 manuscripts in 372 families. Some are formally checked in Lean; OpenAI warns that some of the unformalized results “could have issues.” The most-cited explanation of the problem came from a recent paper’s title: verification abundance, adjudication scarcity. A computer can check that a proof is valid. Deciding whether it proves the right statement, whether it matters, and what it teaches still takes human experts, and the field has a fixed supply of them.
Software. This is where the data is richest, and it all points the same way:
- Faros AI, tracking teams across low- and high-AI-adoption periods, found teams merging 98% more pull requests while review time rose 91%.
- LinearB, analysing 8.1 million pull requests across 4,800 organisations, found AI-generated changes wait 4.6 times longer for review to begin, and are accepted 32.7% of the time against 84.4% for human-written ones.
- A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review at all before merging or closing. Faros saw merges with zero review rise 31.3% in high-adoption periods.
Several of these sources sell code-review tools, so read the exact numbers with care. The direction is consistent across all of them: code got cheaper to write and more expensive to trust.
Professional workflows. OpenAI’s new partnership with the contract-software company Ironclad trained GPT-6 Astra on real contracting workflows. On 11 tasks, Astra met 55% of the evaluation criteria on average. That’s a large improvement over the previous model. It also means someone has to find the other 45% — the missed approval rule, the wrong jurisdiction clause — before the work can be used.
As an affiliate, we earn on qualifying purchases.
Why checking doesn’t get cheaper
It’s tempting to assume AI will simply check AI. Partly, it will. But there are three reasons the human part of verification doesn’t shrink as fast as generation does.
You can’t see the author’s intent. When a colleague writes code or a proof, a reviewer reads their reasoning, asks them questions, and knows their habits. Machine output arrives without any of that. Reviewers report AI-generated code is more taxing to check for exactly this reason: it looks locally clean and gives no clue where its mistakes are.
Formal checking verifies the answer, not the question. A theorem prover confirms a proof proves its stated theorem. Tests confirm code does what the tests check. Neither confirms the statement or the tests capture what was actually needed. That gap is where August’s disputed mathematical “counterexample” failed, and where most real-world software bugs live.
Somebody has to be accountable. This is the deepest reason, and it’s social rather than technical. A contract is signed by a person. A bridge design is stamped by an engineer. A mathematical paper has authors who answer questions at conferences. Responsibility is a legal and institutional function, and an institution can’t hold a model accountable.
As an affiliate, we earn on qualifying purchases.
What happens when referees run out
When checking capacity falls behind output, organisations don’t stop. They degrade, in three predictable ways, and all three are already visible.
Rubber-stamping. The 61% of AI-agent pull requests with no review is the clearest example: output ships because nobody had time to look.
Triage by suspicion. LinearB found 38% of reviewers now deliberately deprioritise AI-generated changes. That’s rational for each reviewer and wasteful overall: good machine work waits behind the assumption that it’s probably bad.
Selection by the producer. OpenAI chose which 722 manuscripts were significant enough to publish out of roughly 4,000 problems. When the referees can’t keep up, the producer’s own filter becomes the de facto review. That’s fine when the producer is careful, and invisible when it isn’t.
As an affiliate, we earn on qualifying purchases.
The apprenticeship paradox
Here’s the part a post-labor analysis can’t skip.
Senior reviewers are made, not hired. A good code reviewer learned by writing code for years. A mathematician who can referee a proof learned by proving things. A contract lawyer who spots the missing clause drafted hundreds of contracts first.
The work AI is absorbing is exactly the work that trained the reviewers. Junior developers ship AI-written pull requests instead of writing code. Junior associates review AI drafts instead of drafting. If that continues, the supply of people capable of adjudicating falls just as demand for adjudication rises.
That’s not a reason to stop using AI. It’s a reason to deliberately protect the path that produces judgement, because nothing else in the economy currently does.
As an affiliate, we earn on qualifying purchases.
The referee premium
Economically, the consequence is straightforward. When one input becomes abundant and its complement stays scarce, value moves to the complement.
Expect a referee premium. The people who can say “this is right” and be held responsible for it — senior engineers, auditors, specialist lawyers, reviewing scientists, safety assessors — become the binding constraint on how much AI output an organisation can actually use. Their time becomes more valuable, not less.
This also explains a pattern this publication has measured directly. In my own AI stack, the metric that matters isn’t cost per token or cost per task. It’s cost per accepted result: model cost plus review and rework, divided by what someone actually uses. In the illustrative example I used — $1 of model time plus four minutes of review at $45 an hour — halving the model price saves 12.5%, and one extra minute of review wipes that out. In that example, review is three-quarters of the bill.
What to do about it
Price verification explicitly. Budget review time as a real cost line next to model spend. If your AI rollout doesn’t measure review hours, you’re measuring the cheap half of the work.
Make machines check what machines can. Formal proof checkers, type systems, test suites, and policy engines turn some judgement into mechanical checking. Every check you formalise frees human reviewers for the parts only they can do.
Tier the review. Not every output needs a senior expert. Route routine work through automated checks and sampling; reserve human attention for high-consequence decisions.
Fund the referees. The mathematics advisory group at the Institute for Advanced Study asked AI labs to fund human understanding of the results they release. That principle generalises: whoever profits from cheap generation should help pay for the verification it requires.
Protect the apprenticeship. Keep some production work deliberately human for people who are learning. It looks inefficient. It’s how you get reviewers in five years.
The take
The first wave of automation anxiety asked which jobs AI would do. The more useful question is which jobs AI makes more necessary.
The answer emerging from mathematics, software and professional workflows is the same: the jobs that check, adjudicate, and take responsibility. Production is becoming cheap and plentiful. Trust is not.
For a post-labor economy, that’s both a warning and an opportunity. The warning is that we are automating the work that trains the people we’ll need most. The opportunity is that accountability — the willingness and ability to stand behind a result — may be the most durable form of human work there is.
The machines can now produce faster than we can check. The organisations, and the economies, that invest in checking will be the ones that actually capture the value.
Sources: OpenAI’s mathematics release (722 manuscripts, 372 families, ~4,000 problems, ~3 hours of Pro compute per result) and the Alon–Bloom–Gowers–Litt–Sawin verification of the Erdős unit-distance counterexample, as covered in this publication’s earlier analysis; “Verification abundance, adjudication scarcity” (arXiv:2608.28997); OpenAI, “Advancing computer use with Ironclad” (6 October 2026: 11 tasks, GPT-6 Astra 55.0% mean rubric score); Faros AI two-year dataset (98% more pull requests, 91% longer review time, 31.3% rise in zero-review merges); LinearB 2026 analysis of 8.1M pull requests across 4,800 organisations (4.6× review wait, 32.7% vs 84.4% acceptance, 38% of reviewers deprioritising AI PRs); Duma et al., “These Aren’t the Reviews You’re Looking For”, EASE 2026 (61% of AI-agent PRs unreviewed) — as reported in secondary coverage. Several code-review data sources are vendors of review tools. The $4 review-cost example is illustrative, from the author’s September 2026 stack analysis. Advisory Group on Mathematics and AI recommendations (29 September 2026). Analysis and framing are the author’s.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
