🔍 Read the full analysis: Ironclad Users: What To Know About OpenAI Training Agents In Software on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training its GPT-6 Astra model on 11 legal, commercial and procurement tasks in hosted copies of contract-management software from Ironclad. Astra met an average 55% of task rubric criteria, while estimated task times were simulated rather than measured customer savings. OpenAI says it is inviting a small number of other software companies to take part in similar work.
OpenAI said on October 6 that it trained its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The results point to a way of teaching AI agents specialized business workflows, but Astra met an average of 55% of evaluation criteria, and OpenAI says human oversight remains necessary.
Ironclad staff and OpenAI employees who use the product selected 11 tasks, including creating nondisclosure agreements, setting up procurement approval processes and updating reusable contract clauses to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes per task. Evaluators scored each task against a rubric containing between eight and 50 criteria, depending on its complexity.
Ironclad provided hosted product environments where models could practise. OpenAI said it created synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. It said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average of 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. It estimated Astra’s time per attempt at 19.2 minutes, versus 37.0 minutes for Sol. An internal model used in Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of criteria on one showcased task.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The results matter because contract and procurement software does more than fill in forms: it applies business rules and approval controls. A workflow may require Finance approval above a spending threshold, Security review for certain requests and Legal review of nonstandard terms. Missing one of those conditions can undermine the process even if an agent completes most of the other steps correctly.
OpenAI’s 55% figure is the average share of evaluation criteria met, not the share of tasks completed. It does not show that Astra reliably completed 55% of workflows, nor that its output is ready for unsupervised use. OpenAI’s post says that losing track of a business rule can limit what an agent can safely be asked to do, and that human oversight still matters.
The time estimates also do not establish a customer productivity gain. OpenAI described them as simulated estimates based on assumed processing and generation speeds, not measured reductions in work time. They apply to the 11 research tasks, not all Ironclad workflows. The figures suggest a research direction, but do not show that customers can currently get correct contract work faster using the agent.
For software vendors, the project offers a potential path to improve agents on difficult tasks in their products. It also raises a strategic issue: if customers increasingly give instructions to an agent rather than use software screens directly, a vendor’s value may depend more on its rules, data model, audit trail and controls than on its interface. OpenAI presents the work as a reason a full contracting platform remains necessary; that is the company’s framing, not a demonstrated market outcome.
contract management software AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
The October 6 post discussed two OpenAI developments. One involved 722 mathematics manuscripts and received wider attention, according to the source material; the Ironclad work received less notice. The name refers to Ironclad, a contract-management software company, rather than a newly announced hardened agent framework.
The project tests a different approach from evaluating an agent on general computer-use tasks: train and assess it inside a specialized product, against workflows and requirements selected by people familiar with the work. The goal OpenAI described was to teach models to understand business rules, carry out multi-step tasks in specialized software and check whether the result meets the initial requirements.
OpenAI said it is inviting a small number of software companies to partner on tasks that current agents cannot reliably complete. It asked prospective partners to bring concrete examples of failures, subject-matter experts, a secure testing environment and data suitable for research. The post does not announce a broad program or identify additional partners.
legal workflow automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Reported Scores Leave Open
The post does not establish how Astra would perform across all Ironclad workflows, with live customer information or under ordinary production conditions. The reported evaluation covers 11 selected research tasks. The source material does not provide enough detail to determine how the criteria were weighted, how often specific types of errors occurred or how the system performed across repeated trials.
It is also unclear whether the simulated task-time estimates would translate into real time savings once people check the work and correct errors. A higher average score does not show whether Astra missed low-impact requirements or failed a safeguard essential to a particular workflow. OpenAI’s reported results do not establish that the system is ready for autonomous contract or procurement decisions.
The company has not named other software partners or described a timetable for future collaborations. The longer-term effect on software vendors—whether agents add value to existing products or reduce customer reliance on their interfaces—remains a strategic possibility, not a result demonstrated by this test.
As an affiliate, we earn on qualifying purchases.
Questions Before Agents Reach Production
OpenAI says it plans to work with a small number of software companies on difficult tasks agents cannot yet complete reliably. Any follow-up should clarify which workflows are being tested, what data can be used and how partners will evaluate results. The announcement provides no date for the next milestone and names no additional participating companies.
Businesses considering agents in contract, finance or customer-record systems can ask vendors for more than a single average score: which criteria were missed, whether mandatory controls failed, how human review works, and how actions are logged and reversed. They should also ask whether reported speed figures come from measured customer use or simulation. The Ironclad results offer an early benchmark for a research effort, not evidence that high-stakes workflows can be handed over without checking.
AI-powered legal document software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company. OpenAI’s October 6 post describes training and testing models on tasks inside hosted copies of its product, not launching a framework called Ironclad.
Did GPT-6 Astra complete 55% of the tasks?
No. OpenAI reported an average of 55% of rubric criteria met across the evaluation. That is not the percentage of tasks completed, and it does not show that the remaining criteria were unimportant.
Are the reported time savings measured in customer deployments?
No. OpenAI said the task times were simulated estimates based on assumed processing and generation speeds. It said they covered the 11 research tasks, not Ironclad workflows generally.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Can businesses use these agents without human review?
The reported results do not support that conclusion. OpenAI said human oversight remains necessary, and the average rubric score does not establish that required approvals or other controls will be followed reliably in production.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
