🔍 Read the full analysis: How Ironclad’s Terms Connect To OpenAI’s Agent Training Inside Software on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI and contract-management company Ironclad tested training a frontier model on tasks inside Ironclad’s software, using synthetic tasks based on public SEC-filed contracts. OpenAI reports GPT-6 Astra met 55% of rubric criteria on average; the reported time estimates are simulated, and the results do not establish that the agent is ready to handle contract work without human review.
OpenAI said Ironclad staff and OpenAI employees selected 11 tasks covering legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.
For the research, Ironclad provided hosted copies of its product for model practice. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. It also said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
OpenAI scored performance against rubrics of 8 to 50 criteria, depending on task complexity. It reported that GPT-6 Astra met 55.0% of criteria on average, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. An internal model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of criteria on one showcase task; that result is for a single example, not the overall average.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Workflow-Level Training Matters
The test shifts attention from whether an agent can operate software in general to whether it can follow company-specific rules while carrying out multi-step work. In contracts and procurement, those rules may govern who approves spending, when Security must review a request and which terms require Legal review. Missing one requirement can undermine the whole workflow, even if the agent completes several other steps correctly.
For software vendors, OpenAI’s invitation to work with a small number of companies signals interest in using realistic product workflows to find and address agent limitations. A vendor could gain evidence about where agents fail and potentially make its product more useful. But improved agents could also become the main way customers interact with a product, making the vendor’s business rules, records and controls more important than its screens.
For customers, the reported score is a reminder that average rubric performance does not establish safe or correct completion. Buyers evaluating agents for consequential work need to know which criteria failed, whether a missed approval or control can stop the process, and who reviews the result. OpenAI’s own description says human oversight remains necessary when agents may lose track of business rules.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
The October 6 post, titled “Advancing computer use with Ironclad,” describes a research collaboration rather than the release of a new agent framework called Ironclad. Ironclad is a contract-management software company; the work tested models on tasks within its product. OpenAI framed the goal as teaching models to understand business rules, perform multi-step work in specialised software and check completed work against the original requirements.
OpenAI also reported estimated time per attempt of 19.2 minutes for GPT-6 Astra, compared with 37.0 minutes for GPT-5.6 Sol. The company’s footnote says those figures are simulated estimates based on assumed processing and generation speeds, not customer-measured time savings. They apply to the 11 research tasks, not to Ironclad workflows generally. They cannot establish that the agent is faster than a person who completes the work correctly.
The reported 55% figure is the average share of criteria met, not the percentage of tasks completed successfully. A rubric score can include partial credit, but the source does not provide enough detail here to determine how each missed criterion affected a task’s practical outcome. OpenAI’s own procurement example underscores the issue: failing to preserve a required Finance, Security or Legal approval could make a workflow unacceptable even if other steps were performed.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Establish
The post does not establish how the models would perform across Ironclad’s full range of customer workflows, on live customer data or in routine commercial use. OpenAI says the evaluation covered 11 selected research tasks; the source material does not describe independent testing or provide a broader deployment record. It is also unclear how often particular high-consequence requirements were missed, or whether some errors would be more serious than others.
The time figures are explicitly simulated rather than observed savings, while the criteria score is not a pass rate. No customer productivity gains, error rates in production or return on investment are confirmed by the reported test. OpenAI’s stated data restrictions describe what it says was used for this research; the source does not provide further details about data governance, retention or any future partner arrangements.
AI-powered contract review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What OpenAI and Buyers May Do Next
OpenAI said it is inviting a small number of software companies to partner on tasks current agents cannot reliably complete. It asked prospective partners to bring a concrete example of a failing task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not name additional partners or set a timetable for future results.
Before deploying agents in contract, finance or customer-record systems, organisations should ask vendors to report which requirements failed, how approval controls are enforced and where a human must intervene. They should also distinguish simulated benchmarks from measured results in their own workflows. Further evaluations, partnership announcements or production data would be needed to show whether this training approach produces reliable operational gains.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They tested training and evaluating models on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.
What does GPT-6 Astra’s 55% score mean?
OpenAI says Astra met 55% of evaluation criteria on average. That is not a claim that it completed 55% of tasks successfully, and it does not show that the remaining requirements were unimportant.
Did the test show customers would save time?
No. OpenAI described the reported 19.2-minute estimate for Astra as simulated, based on assumed processing and generation speeds. It was not a measurement of customer time saved.
What data did OpenAI say it used?
OpenAI said it generated synthetic tasks from publicly filed SEC EDGAR contracts, filtered to remove personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can businesses use the agent for contract work without review?
The reported results do not support that conclusion. OpenAI’s post says human oversight still matters, and the average score does not establish that all required approvals or business rules were preserved.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
