The Ironclad Details Behind OpenAI’s Training Of Agents In Your Software
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Ironclad Details Behind OpenAI’s Training Of Agents In Your Software on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and testing its GPT-6 Astra model in hosted copies of contract-management software from Ironclad, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task rubric criteria; its estimated completion time was simulated, not measured in customer use. OpenAI is inviting a small number of other software companies to explore similar partnerships.

OpenAI said on Oct. 6 that it trained and tested its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The project offers a look at how AI developers may work with software vendors to teach agents specialized business workflows, but Astra met an average of 55% of evaluation criteria and still requires human oversight, according to the company’s account.

OpenAI and Ironclad selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. The tasks were scored against rubrics containing between eight and 50 criteria, depending on complexity.

Ironclad supplied hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data for this work.

OpenAI reported that GPT-6 Astra met an average 55% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. Astra’s estimated time per attempt was 19.2 minutes, versus 37 minutes for Sol. An internal OpenAI model used in Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of criteria. These are OpenAI-reported results, not an independent evaluation.

At a glance
reportWhen: Described by OpenAI on Oct. 6; further…
The developmentOpenAI published details on Oct. 6 of a project training and evaluating a frontier model on workflows inside Ironclad’s contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Business Rules Matter More Than Clicks

The project matters because professional software tasks are not simply sequences of clicks. A contract workflow may depend on spending thresholds, security checks, jurisdiction-specific language and legal review. An agent can appear to complete most of a task while missing a rule that determines whether the result is safe to use.

OpenAI’s 55% figure is the average share of criteria met, not the share of tasks completed successfully. That distinction matters in high-consequence work: a procurement workflow missing a required Finance or Security approval may be unusable even if it satisfies several other criteria. OpenAI’s own discussion recognizes that losing track of a business rule limits what a company can confidently delegate. Human review remains part of the process described.

For software vendors, the partnership model offers a way to improve agents on real, specialized workflows. It also creates a strategic question: if customers increasingly tell an agent what they need rather than working through a product’s screens, the vendor’s value may depend more on its business rules, records, audit trail and controls than on its interface. That is an implication of the development, not a reported change in how Ironclad customers currently work.

Amazon

contract management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became a Test Environment

The OpenAI post was titled “Advancing computer use with Ironclad.” ThorstenMeyerAI.com said some AI news trackers interpreted “Ironclad” as the name of an agent framework; in this case, it refers to the contract-management software company. The post describes work inside Ironclad’s product rather than the launch of a new framework.

The project joins three elements: tasks chosen by people familiar with the software and workflows, a hosted product environment for model practice, and rubrics intended to measure whether the agent followed task requirements. OpenAI says GPT-6 Astra is the first frontier model it trained this way. Its account describes a research effort, not a general product release or evidence that the model is ready to handle contracting work without checks.

OpenAI also said it is looking for a small number of software-company partners. It asked prospective partners to bring a concrete example of a task current agents cannot reliably complete, people with deep knowledge of the work, a secure test environment and data that can be used safely for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Cannot Show Yet

The figures do not show how often Astra would produce fully acceptable results in routine customer use. The 55% score averages criteria met across the evaluation; the source material does not provide the full results for every task or explain which specific requirements were most often missed. A high score on one showcase task does not establish reliable performance across the other workflows.

The time estimates also are not observed productivity gains. OpenAI’s footnote says they are simulated estimates based on assumed processing and generation speeds, covering the 11 research tasks rather than Ironclad workflows generally. The source does not report customer deployments, measured time saved, or a comparison of fully reviewed agent output with work completed by experienced professionals.

OpenAI says it did not use specified customer or internal contract data, but details of the data filtering and evaluation process are not independently established in the source material. It is also unclear which software companies will participate next, what tasks they will test, or whether later work will use the same scoring approach.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Tests for Software Agents

OpenAI says it plans to work with a small number of software companies on tasks agents still struggle to complete reliably. The requested ingredients—specific failure cases, expert input, secure environments and suitable research data—suggest that future projects will depend on vendors defining both the workflow and the rules that count as success.

For buyers, the immediate practical issue is not whether an agent can perform steps in an interface, but whether it can meet every requirement and make missed requirements visible to a reviewer. The results reported so far do not establish that an agent can replace experienced staff on contract or procurement work. Further partner projects, fuller task-level results and measured performance in real use would help show whether the approach produces dependable outcomes.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this OpenAI project?

Ironclad is a contract-management software company. OpenAI’s post describes training and testing a model using hosted copies of Ironclad’s product, not a framework named Ironclad.

What does Astra’s 55% score mean?

It is the average share of evaluation criteria met across the tasks, not the percentage of tasks completed successfully. The score does not indicate that the work is ready to use without review.

Did OpenAI measure customer time savings?

No such measured customer savings are reported in the source material. OpenAI described the completion times as simulated estimates based on assumed processing and generation speeds.

What data did OpenAI say it used?

OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

Can Astra handle contract workflows without human review?

The reported results do not support that conclusion. OpenAI’s account says agents can lose track of business rules and that human oversight still matters; the average criteria score was 55%.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Samsung chip workers will get an average $340k bonus as AI profits soar

Samsung plans an average bonus of $340,000 for its chip workers as profits from AI-related semiconductor sales surge, highlighting industry gains.

Total Kills Over/Under 60.5 In Game 2?

A new betting market on Polymarket for Game 2’s total kills over/under 60.5 has been listed, sparking increased betting interest amid ongoing esports events.

Majority of Americans Support Ban on Surveillance Pricing and Electronic Shelf Labels

A new survey shows 68% of Americans oppose surveillance pricing and electronic shelf labels, citing concerns over rising grocery costs and privacy.

Spirit Airlines Spent $1.61 For Every $1 It Took In — New Filing Shows Why It Couldn’t Be Saved

A new bankruptcy filing shows Spirit Airlines spent $1.61 for every dollar earned in March, highlighting its financial decline and why a bailout was unlikely.