🔍 Read the full analysis: OpenAI Is Training Agents In Software. Here Are The Ironclad Details To Check on ThorstenMeyerAI.com
Get smart everyday buys delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training a frontier model in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while its estimated completion times were simulated rather than measured customer savings.
OpenAI said on October 6 that it trained and evaluated a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. The results offer an early measure of how well an agent can follow rules inside specialized business software, but the reported averages fall well short of showing that the work can be trusted without human review.
Ironclad staff and OpenAI employees who use the product selected 11 workflow tasks, including setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to match a selected jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Depending on complexity, the tasks were assessed against 8 to 50 criteria.
Ironclad provided hosted product environments in which models could practise. OpenAI said it created synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average 55.0% of the criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used in Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of criteria. These are rubric scores across the tasks—not the share of tasks completed successfully. OpenAI also estimated 19.2 minutes per Astra attempt, versus 37 minutes for GPT-5.6 Sol.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Workflow Accuracy Matters
The work tests whether AI agents can do more than operate buttons and fields: they must retain a company’s rules across a multi-step process and produce a result that meets its requirements. In contract and procurement work, missing one approval condition can matter more than getting several routine steps right. A score of 55% of criteria met should not be read as a workflow that is 55% safe or useful.
OpenAI’s own example describes a procurement process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. If an agent omits one of those controls, the process may route a purchase incorrectly. That makes the result relevant to buyers considering agents for contracts, finance or customer records: performance must be checked against specific requirements, not inferred from a single average score.
For software vendors, the project points to a possible shift in how agents are developed: vendors can provide realistic environments and expert-defined tasks for model training and evaluation. Better agents could make a product more useful, but they could also become the main way customers interact with it. In that case, the vendor’s lasting value may depend on its business rules, data, audit trail and controls, rather than only its interface.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Built
The October 6 post was titled “Advancing computer use with Ironclad.” Its subject is Ironclad, a contract-management software company, not a newly announced general-purpose agent framework. OpenAI framed the research around training models to understand business rules, carry out multi-step work in specialized software and check the result against the original requirements.
The project’s test set was small and defined jointly by people familiar with the product and its workflows. It used 11 selected tasks and criteria tailored to each task, with different numbers of checks depending on complexity. OpenAI’s description presents the work as a research exercise, not a broad customer rollout or a claim that agents can now handle contract work independently.
OpenAI also invited a small number of software companies to discuss similar partnerships. It said prospective partners should bring concrete examples of tasks agents cannot reliably complete, experts who understand the work, a secure test environment and data that can safely be used for research. The post says a full contracting platform remains important because agents need to preserve the controls on which teams rely.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The 19.2-minute estimate is not a timed result from customer use. OpenAI says the figures were simulated using assumed processing and generation speeds and apply to the 11 research tasks, not Ironclad workflows generally. The source material does not provide measured customer time savings or evidence of deployment at scale.
The average 55% score also does not disclose, in the figures provided, which specific criteria Astra missed on each task or how often an error would create a material business risk. A high result on one showcase task does not establish consistent performance across other workflows. OpenAI says human oversight remains necessary, but the materials cited here do not specify a product release plan, operational safeguards or a threshold for reducing that oversight.
The source describes OpenAI’s account of its data use and testing. It does not provide independent verification of those statements or details about whether additional companies have agreed to participate. The eventual effect on software vendors’ customer relationships and business models remains uncertain.
As an affiliate, we earn on qualifying purchases.
What Partners May Test Next
OpenAI says it is seeking a small number of software-company partners to work on tasks current agents cannot reliably complete. The proposed partners would supply concrete failure cases, knowledgeable staff, secure testing environments and research-safe data. The post does not name further participants or give a timetable for new evaluations.
For future work to show progress clearly, readers will need task-by-task results, details of missed criteria and evidence from measured use—not simulated time estimates alone. Until those details are available, the Ironclad project is best read as an early test of agent performance in specialized software, with human review still part of the process.
procurement approval process software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company. OpenAI’s post describes training and evaluating models in hosted copies of its product, not announcing a new agent framework called Ironclad.
What does GPT-6 Astra’s 55% score mean?
It is the average share of rubric criteria met across the selected tasks. It does not mean Astra completed 55% of the tasks, and it does not establish that a workflow is safe to use without review.
Did the reported time estimate show customer time savings?
No. OpenAI said its 19.2-minute estimate was simulated from assumed processing and generation speeds. It was not measured in customer use and covered the 11 research tasks, not Ironclad workflows generally.
What data did OpenAI say it used?
OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can companies use these agents for contract work without human oversight?
The reported results do not support that conclusion. OpenAI’s post says human oversight remains necessary, and the average criteria score leaves substantial room for errors or missed requirements.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
