🔍 Read the full analysis: The Ironclad Details Behind OpenAI’s Training Of Agents In Your Software on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training and testing its GPT-6 Astra model in hosted copies of contract-management software from Ironclad, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task rubric criteria; its estimated completion time was simulated, not measured in customer use. OpenAI is inviting a small number of other software companies to explore similar partnerships.
OpenAI said on Oct. 6 that it trained and tested its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The project offers a look at how AI developers may work with software vendors to teach agents specialized business workflows, but Astra met an average of 55% of evaluation criteria and still requires human oversight, according to the company’s account.
OpenAI and Ironclad selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. The tasks were scored against rubrics containing between eight and 50 criteria, depending on complexity.
Ironclad supplied hosted copies of its product for model practice. OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data for this work.
OpenAI reported that GPT-6 Astra met an average 55% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. Astra’s estimated time per attempt was 19.2 minutes, versus 37 minutes for Sol. An internal OpenAI model used in Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of criteria. These are OpenAI-reported results, not an independent evaluation.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Business Rules Matter More Than Clicks
The project matters because professional software tasks are not simply sequences of clicks. A contract workflow may depend on spending thresholds, security checks, jurisdiction-specific language and legal review. An agent can appear to complete most of a task while missing a rule that determines whether the result is safe to use.
OpenAI’s 55% figure is the average share of criteria met, not the share of tasks completed successfully. That distinction matters in high-consequence work: a procurement workflow missing a required Finance or Security approval may be unusable even if it satisfies several other criteria. OpenAI’s own discussion recognizes that losing track of a business rule limits what a company can confidently delegate. Human review remains part of the process described.
For software vendors, the partnership model offers a way to improve agents on real, specialized workflows. It also creates a strategic question: if customers increasingly tell an agent what they need rather than working through a product’s screens, the vendor’s value may depend more on its business rules, records, audit trail and controls than on its interface. That is an implication of the development, not a reported change in how Ironclad customers currently work.
contract management software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Ironclad Became a Test Environment
The OpenAI post was titled “Advancing computer use with Ironclad.” ThorstenMeyerAI.com said some AI news trackers interpreted “Ironclad” as the name of an agent framework; in this case, it refers to the contract-management software company. The post describes work inside Ironclad’s product rather than the launch of a new framework.
The project joins three elements: tasks chosen by people familiar with the software and workflows, a hosted product environment for model practice, and rubrics intended to measure whether the agent followed task requirements. OpenAI says GPT-6 Astra is the first frontier model it trained this way. Its account describes a research effort, not a general product release or evidence that the model is ready to handle contracting work without checks.
OpenAI also said it is looking for a small number of software-company partners. It asked prospective partners to bring a concrete example of a task current agents cannot reliably complete, people with deep knowledge of the work, a secure test environment and data that can be used safely for research.
As an affiliate, we earn on qualifying purchases.
What the Evaluation Cannot Show Yet
The figures do not show how often Astra would produce fully acceptable results in routine customer use. The 55% score averages criteria met across the evaluation; the source material does not provide the full results for every task or explain which specific requirements were most often missed. A high score on one showcase task does not establish reliable performance across the other workflows.
The time estimates also are not observed productivity gains. OpenAI’s footnote says they are simulated estimates based on assumed processing and generation speeds, covering the 11 research tasks rather than Ironclad workflows generally. The source does not report customer deployments, measured time saved, or a comparison of fully reviewed agent output with work completed by experienced professionals.
OpenAI says it did not use specified customer or internal contract data, but details of the data filtering and evaluation process are not independently established in the source material. It is also unclear which software companies will participate next, what tasks they will test, or whether later work will use the same scoring approach.
AI-powered contract review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Next Tests for Software Agents
OpenAI says it plans to work with a small number of software companies on tasks agents still struggle to complete reliably. The requested ingredients—specific failure cases, expert input, secure environments and suitable research data—suggest that future projects will depend on vendors defining both the workflow and the rules that count as success.
For buyers, the immediate practical issue is not whether an agent can perform steps in an interface, but whether it can meet every requirement and make missed requirements visible to a reviewer. The results reported so far do not establish that an agent can replace experienced staff on contract or procurement work. Further partner projects, fuller task-level results and measured performance in real use would help show whether the approach produces dependable outcomes.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this OpenAI project?
Ironclad is a contract-management software company. OpenAI’s post describes training and testing a model using hosted copies of Ironclad’s product, not a framework named Ironclad.
What does Astra’s 55% score mean?
It is the average share of evaluation criteria met across the tasks, not the percentage of tasks completed successfully. The score does not indicate that the work is ready to use without review.
Did OpenAI measure customer time savings?
No such measured customer savings are reported in the source material. OpenAI described the completion times as simulated estimates based on assumed processing and generation speeds.
What data did OpenAI say it used?
OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
Can Astra handle contract workflows without human review?
The reported results do not support that conclusion. OpenAI’s account says agents can lose track of business rules and that human oversight still matters; the average criteria score was 55%.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
