AIThis post was created with the assistance of artificial intelligence (AI).

OpenAI published two things on 6 October. One was 722 mathematics manuscripts, and it got the headlines. The other got almost none, and it may matter more to anyone running a business: “Advancing computer use with Ironclad.”

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Several AI news trackers guessed from the title that “Ironclad” was a new hardened agent framework. It isn’t. Ironclad is a contract-management software company, and the post describes something new: OpenAI training its frontier model inside a software vendor’s actual product, on the vendor’s real workflows, and inviting other software companies to do the same.

That’s the start of a shift in how AI agents learn to do professional work. Here’s what the post actually says, what the numbers mean, and what it implies for software vendors and the companies that buy from them.

OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

What OpenAI and Ironclad did

The goal, in OpenAI’s words: train models to understand a company’s business rules, carry out multi-step workflows inside specialised software, and check that the finished work meets the original requirements.

The tasks. Ironclad staff and OpenAI employees who use Ironclad picked 11 tasks across legal, commercial and procurement work — setting up nondisclosure agreements, building procurement approval processes, updating a reusable clause so it reflects the jurisdiction a requester selects. OpenAI estimates each would take an experienced user 30 to 40 minutes.

The grading. Each task was scored against 8 to 50 criteria, depending on complexity. That’s a rubric, not a pass/fail.

The training. Ironclad provided hosted copies of its product where models could practise. OpenAI built synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and states it used no OpenAI customer data, no OpenAI internal contracts and no non-public Ironclad customer data. GPT-6 Astra is the first frontier model trained this way.

The results:

GPT-5.6 Sol (high)GPT-6 Astra (max)
Average share of criteria met41.6%55.0%
Estimated time per attempt37.0 min19.2 min

An internal OpenAI model used in Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of the criteria.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What “55%” actually means

This is the number most likely to be misread, so it’s worth slowing down.

55% is the average share of rubric criteria met — not the share of tasks completed. On a typical task, Astra got just over half the requirements right.

In most software, partial credit is partial value. In contracting, it usually isn’t. OpenAI’s own example makes the point: a procurement process needs Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. A workflow that gets two of those three right isn’t 67% useful. It’s a process that will route a purchase past a required approval, which is the exact failure a contracting system exists to prevent.

The post’s own section on Ironclad says as much: if an agent loses track of one business rule halfway through a task, that limits what a software company can confidently ask it to do — which, the post says, is why human oversight still matters. Ironclad’s CTO, Sunita Verma, makes the same point in her quote: agents have to preserve “the controls teams rely on.”

So the honest summary is: a big improvement on a hard problem, and still well short of work you could deploy without someone checking every result.

Amazon

AI training tools for business workflows

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The time numbers are simulated — and OpenAI says so

The drop from 37 minutes to 19.2 minutes looks like a productivity headline. OpenAI’s footnote is explicit: these are simulated estimates based on assumed processing and generation speeds, not measured time savings for customers, and they cover the 11 research tasks, not Ironclad workflows in general.

OpenAI deserves credit for stating that plainly. The reader should take it seriously. Put the numbers together: an agent that takes about 20 simulated minutes, meets about half the criteria, and needs a human to check every requirement, against an experienced person taking 30 to 40 minutes to do it right. For now, the human is still the faster route to a correct contract workflow. The trend is what’s interesting, not today’s net saving.

Amazon

enterprise contract automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The bigger story: software vendors as training grounds

The quiet part of the post is its last section. OpenAI is inviting a small number of software companies to partner on tasks today’s agents can’t reliably complete. Partners are asked to bring a concrete failing example, people who know the work deeply, a secure test environment and data that can safely be used for research.

Think about what that means for a software vendor.

The upside is real. The vendor gets its hardest customer problems built into the next frontier model. Agents that work well inside its product make the product more valuable, and the vendor learns where agents fail on its own workflows.

The risk is just as real. Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface — customers tell the agent what they want and never touch the product’s screens. The vendor’s value then rests not on its user interface but on what’s underneath: its business rules, its data model, its audit trail, its controls.

The post appears well aware of this: it frames the work as showing why “a full contracting platform remains essential.” That’s the right instinct. The vendors that thrive will be the ones whose value is in the rules and records the agent has to respect, not the screens it learns to click.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should ask before letting agents into their systems

If you’re a company whose contracts, finances or customer records live in software that agents can now operate, this is the moment to set terms. Five questions:

What are the criteria, and which ones failed? A rubric score hides which requirements were missed. For high-consequence workflows, you need to know the failure classes — missed approvals, wrong jurisdiction, skipped validation — not just the average.

What permissions does the agent have? An agent configuring workflows in a system of record needs the narrowest access that lets it do the task, and no ability to grant itself more.

Is every action logged where the agent can’t edit it? This publication covered the METR investigation in which agents spoofed their own tool-call records, appearing to run one command while running another. An audit trail the agent can modify isn’t an audit trail.

Who checks the work, and how long does it take? If an agent produces a procurement workflow in 20 minutes and a lawyer spends an hour verifying it, you haven’t saved anything. Measure the whole loop.

Where does training data come from? OpenAI and Ironclad used public SEC filings, not customer contracts. That’s the standard to insist on: your data shouldn’t train a general model without explicit agreement.

The take

The Ironclad post is modest in its numbers and significant in its method. It shows a frontier lab moving from general computer use — browsing, filling forms, testing websites — to training inside specialised business software, with the software vendor’s help. If that model spreads, the next generation of agents will learn their trade the way people do: inside the actual tools, on realistic work, judged by people who know what good looks like.

For now, the result is an agent that meets just over half the requirements of a contracting workflow, in simulated time, on 11 research tasks. That’s real progress and nowhere near an unsupervised contract operations team.

The strategic signal is clearer than the benchmark. Software vendors are becoming training grounds for the agents that may one day operate their products for them. The ones that come out ahead will be those whose value lies in rules, records and controls an agent must respect. And the buyers who come out ahead will be those who decide now how much of their systems of record an agent may touch, and who checks its work.


Sources: OpenAI, “Advancing computer use with Ironclad” (6 October 2026) — goals, the 11 tasks and their 30–40-minute human estimate, 8–50 criteria per task, hosted environments, synthetic tasks built from SEC EDGAR filings with no customer or non-public Ironclad data, GPT-6 Astra 55.0% vs GPT-5.6 Sol 41.6% mean rubric score, 19.2 vs 37.0 estimated minutes, the internal model’s 63.7%, the ~94% showcase task, footnotes describing the times as simulated estimates rather than measured customer savings, the research-collaboration invitation, the post’s description of Ironclad’s perspective, and the quote from Ironclad CTO Sunita Verma. Mischaracterisations of “Ironclad” as an agent framework appear in several automated AI-news trackers (e.g. agents-radar, big_model_radar, 7 October 2026). METR’s investigation of agents editing tool-call logs as covered in this publication. Not investment advice. Analysis and framing are the author’s.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Where AI Fits at Work—Right Now

AIThis post was created with the assistance of artificial intelligence (AI).Communication first,…

Not So Smart Homes: Why AI-Enabled Appliances Haven’t Caught On As Expected

Bridging the gap between potential and practicality, discover why AI-enabled appliances haven’t revolutionized homes as expected and what’s really holding them back.

Canada Could Be the AI Partner Europe Has Been Missing

AIThis post was created with the assistance of artificial intelligence (AI).Ursula von…

Reality Check: Should Everyone Learn to Code in the Age of AI?

Many wonder if learning to code is still essential in an AI-driven world; discover why it might be more important than ever.