OpenAI Is Training Agents In Software. Here Are The Ironclad Details To Check
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Is Training Agents In Software. Here Are The Ironclad Details To Check on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get smart everyday buys delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in hosted copies of contract-management software Ironclad, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while its estimated completion times were simulated rather than measured customer savings.

OpenAI said on October 6 that it trained and evaluated a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. The results offer an early measure of how well an agent can follow rules inside specialized business software, but the reported averages fall well short of showing that the work can be trusted without human review.

Ironclad staff and OpenAI employees who use the product selected 11 workflow tasks, including setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to match a selected jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Depending on complexity, the tasks were assessed against 8 to 50 criteria.

Ironclad provided hosted product environments in which models could practise. OpenAI said it created synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

OpenAI reported that GPT-6 Astra met an average 55.0% of the criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used in Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of criteria. These are rubric scores across the tasks—not the share of tasks completed successfully. OpenAI also estimated 19.2 minutes per Astra attempt, versus 37 minutes for GPT-5.6 Sol.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of a partnership with Ironclad to train and evaluate models on specialized contract-management workflows.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Workflow Accuracy Matters

The work tests whether AI agents can do more than operate buttons and fields: they must retain a company’s rules across a multi-step process and produce a result that meets its requirements. In contract and procurement work, missing one approval condition can matter more than getting several routine steps right. A score of 55% of criteria met should not be read as a workflow that is 55% safe or useful.

OpenAI’s own example describes a procurement process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. If an agent omits one of those controls, the process may route a purchase incorrectly. That makes the result relevant to buyers considering agents for contracts, finance or customer records: performance must be checked against specific requirements, not inferred from a single average score.

For software vendors, the project points to a possible shift in how agents are developed: vendors can provide realistic environments and expert-defined tasks for model training and evaluation. Better agents could make a product more useful, but they could also become the main way customers interact with it. In that case, the vendor’s lasting value may depend on its business rules, data, audit trail and controls, rather than only its interface.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Built

The October 6 post was titled “Advancing computer use with Ironclad.” Its subject is Ironclad, a contract-management software company, not a newly announced general-purpose agent framework. OpenAI framed the research around training models to understand business rules, carry out multi-step work in specialized software and check the result against the original requirements.

The project’s test set was small and defined jointly by people familiar with the product and its workflows. It used 11 selected tasks and criteria tailored to each task, with different numbers of checks depending on complexity. OpenAI’s description presents the work as a research exercise, not a broad customer rollout or a claim that agents can now handle contract work independently.

OpenAI also invited a small number of software companies to discuss similar partnerships. It said prospective partners should bring concrete examples of tasks agents cannot reliably complete, experts who understand the work, a secure test environment and data that can safely be used for research. The post says a full contracting platform remains important because agents need to preserve the controls on which teams rely.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Show

The 19.2-minute estimate is not a timed result from customer use. OpenAI says the figures were simulated using assumed processing and generation speeds and apply to the 11 research tasks, not Ironclad workflows generally. The source material does not provide measured customer time savings or evidence of deployment at scale.

The average 55% score also does not disclose, in the figures provided, which specific criteria Astra missed on each task or how often an error would create a material business risk. A high result on one showcase task does not establish consistent performance across other workflows. OpenAI says human oversight remains necessary, but the materials cited here do not specify a product release plan, operational safeguards or a threshold for reducing that oversight.

The source describes OpenAI’s account of its data use and testing. It does not provide independent verification of those statements or details about whether additional companies have agreed to participate. The eventual effect on software vendors’ customer relationships and business models remains uncertain.

Amazon

AI-powered NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Partners May Test Next

OpenAI says it is seeking a small number of software-company partners to work on tasks current agents cannot reliably complete. The proposed partners would supply concrete failure cases, knowledgeable staff, secure testing environments and research-safe data. The post does not name further participants or give a timetable for new evaluations.

For future work to show progress clearly, readers will need task-by-task results, details of missed criteria and evidence from measured use—not simulated time estimates alone. Until those details are available, the Ironclad project is best read as an early test of agent performance in specialized software, with human review still part of the process.

Amazon

procurement approval process software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training and evaluating models in hosted copies of its product, not announcing a new agent framework called Ironclad.

What does GPT-6 Astra’s 55% score mean?

It is the average share of rubric criteria met across the selected tasks. It does not mean Astra completed 55% of the tasks, and it does not establish that a workflow is safe to use without review.

Did the reported time estimate show customer time savings?

No. OpenAI said its 19.2-minute estimate was simulated from assumed processing and generation speeds. It was not measured in customer use and covered the 11 research tasks, not Ironclad workflows generally.

What data did OpenAI say it used?

OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can companies use these agents for contract work without human oversight?

The reported results do not support that conclusion. OpenAI’s post says human oversight remains necessary, and the average criteria score leaves substantial room for errors or missed requirements.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

Europe pushes for reliable AI access, sovereignty, and safety at the Évian G7 summit, challenging U.S. control over frontier models and global AI governance.

Software engineering. The canonical case.

New data shows junior developer hiring dropped 40% since 2022, while senior engineers see augmentation. The sector reveals heterogeneous impacts of AI.

Outcome-First Decisions: The Friction Is The Feature

A new decision framework prioritizes testing and evidence over plans, aiming to reduce costly mistakes and improve decision accuracy.

The labor share. Is value really moving from labor to capital? The data isn’t on anyone’s side yet.

Examining whether AI is shifting value from labor to capital, the data shows stable aggregate labor share but rising marginal signals. What does this mean?