AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How Ironclad’s Terms Connect To OpenAI’s Agent Training Inside Software on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI and contract-management company Ironclad tested training a frontier model on tasks inside Ironclad’s software, using synthetic tasks based on public SEC-filed contracts. OpenAI reports GPT-6 Astra met 55% of rubric criteria on average; the reported time estimates are simulated, and the results do not establish that the agent is ready to handle contract work without human review.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on contract and procurement tasks inside hosted copies of Ironclad’s software, reporting an average score of 55% of evaluation criteria met. The work matters because it tests a way to train AI agents on the workflows of specialised business software, but OpenAI’s reported results do not show that the model can complete those tasks reliably without human oversight.

OpenAI said Ironclad staff and OpenAI employees selected 11 tasks covering legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

For the research, Ironclad provided hosted copies of its product for model practice. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. It also said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

OpenAI scored performance against rubrics of 8 to 50 criteria, depending on task complexity. It reported that GPT-6 Astra met 55.0% of criteria on average, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. An internal model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of criteria on one showcase task; that result is for a single example, not the overall average.

At a glance
reportWhen: Published October 6; further partnershi…
The developmentAn OpenAI post dated October 6 describes training and evaluating a frontier model on professional workflows in Ironclad’s contract-management software and invites other software firms to explore similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Workflow-Level Training Matters

The test shifts attention from whether an agent can operate software in general to whether it can follow company-specific rules while carrying out multi-step work. In contracts and procurement, those rules may govern who approves spending, when Security must review a request and which terms require Legal review. Missing one requirement can undermine the whole workflow, even if the agent completes several other steps correctly.

For software vendors, OpenAI’s invitation to work with a small number of companies signals interest in using realistic product workflows to find and address agent limitations. A vendor could gain evidence about where agents fail and potentially make its product more useful. But improved agents could also become the main way customers interact with a product, making the vendor’s business rules, records and controls more important than its screens.

For customers, the reported score is a reminder that average rubric performance does not establish safe or correct completion. Buyers evaluating agents for consequential work need to know which criteria failed, whether a missed approval or control can stop the process, and who reviews the result. OpenAI’s own description says human oversight remains necessary when agents may lose track of business rules.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

The October 6 post, titled “Advancing computer use with Ironclad,” describes a research collaboration rather than the release of a new agent framework called Ironclad. Ironclad is a contract-management software company; the work tested models on tasks within its product. OpenAI framed the goal as teaching models to understand business rules, perform multi-step work in specialised software and check completed work against the original requirements.

OpenAI also reported estimated time per attempt of 19.2 minutes for GPT-6 Astra, compared with 37.0 minutes for GPT-5.6 Sol. The company’s footnote says those figures are simulated estimates based on assumed processing and generation speeds, not customer-measured time savings. They apply to the 11 research tasks, not to Ironclad workflows generally. They cannot establish that the agent is faster than a person who completes the work correctly.

The reported 55% figure is the average share of criteria met, not the percentage of tasks completed successfully. A rubric score can include partial credit, but the source does not provide enough detail here to determine how each missed criterion affected a task’s practical outcome. OpenAI’s own procurement example underscores the issue: failing to preserve a required Finance, Security or Legal approval could make a workflow unacceptable even if other steps were performed.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Establish

The post does not establish how the models would perform across Ironclad’s full range of customer workflows, on live customer data or in routine commercial use. OpenAI says the evaluation covered 11 selected research tasks; the source material does not describe independent testing or provide a broader deployment record. It is also unclear how often particular high-consequence requirements were missed, or whether some errors would be more serious than others.

The time figures are explicitly simulated rather than observed savings, while the criteria score is not a pass rate. No customer productivity gains, error rates in production or return on investment are confirmed by the reported test. OpenAI’s stated data restrictions describe what it says was used for this research; the source does not provide further details about data governance, retention or any future partner arrangements.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What OpenAI and Buyers May Do Next

OpenAI said it is inviting a small number of software companies to partner on tasks current agents cannot reliably complete. It asked prospective partners to bring a concrete example of a failing task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not name additional partners or set a timetable for future results.

Before deploying agents in contract, finance or customer-record systems, organisations should ask vendors to report which requirements failed, how approval controls are enforced and where a human must intervene. They should also distinguish simulated benchmarks from measured results in their own workflows. Further evaluations, partnership announcements or production data would be needed to show whether this training approach produces reliable operational gains.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested training and evaluating models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.

What does GPT-6 Astra’s 55% score mean?

OpenAI says Astra met 55% of evaluation criteria on average. That is not a claim that it completed 55% of tasks successfully, and it does not show that the remaining requirements were unimportant.

Did the test show customers would save time?

No. OpenAI described the reported 19.2-minute estimate for Astra as simulated, based on assumed processing and generation speeds. It was not a measurement of customer time saved.

What data did OpenAI say it used?

OpenAI said it generated synthetic tasks from publicly filed SEC EDGAR contracts, filtered to remove personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can businesses use the agent for contract work without review?

The reported results do not support that conclusion. OpenAI’s post says human oversight still matters, and the average score does not establish that all required approvals or business rules were preserved.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Scaling AI Responsibly: Gilbert + Tobin’s Partnership With OpenAI

Australian law firm Gilbert + Tobin collaborates with OpenAI to govern and scale AI use, emphasizing controlled deployment amid legal confidentiality concerns.

Border Felony Charges: Disrupting The Pulse Of Trade And Supply-Chain Operations

U.S. authorities have filed felony charges against a citizen for deleting phone data at the border, raising concerns for trade and supply-chain operations.

Court Agrees With EFF: Utah’s VPN Law Demands A Technical Impossibility

A federal judge preliminarily blocked Utah’s VPN provisions in SB 73, finding the law likely burdens people and businesses beyond the state.

Organize Restorative Justice Cases With A Defined Program Workflow

An IdeaNavigator AI proposal outlines a case workflow for restorative justice programs, but calls for a small pilot before testing demand.