🔍 Read the full analysis: OpenAI’s Software Training For Agents Puts Ironclad’s Fine Print In Focus on ThorstenMeyerAI.com
Get the little things that make your day delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training and evaluating a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. The model met an average 55% of task criteria, while estimated completion times were simulated; OpenAI says human oversight remains important.
OpenAI said on October 6 that it had trained and tested a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement workflows. The work offers a look at how AI agents might learn specialized business processes, but the reported model met an average 55% of evaluation criteria—not 55% of tasks—and OpenAI’s time estimates were simulated rather than measured customer savings.
Ironclad staff and OpenAI employees who use the product selected 11 multi-step tasks, including setting up nondisclosure agreements, creating procurement approval processes and revising a reusable contract clause according to a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Depending on the task’s complexity, evaluators scored the work against 8 to 50 criteria.
Ironclad provided hosted product environments where models could practise. OpenAI said it created synthetic tasks using publicly filed contracts from the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data for the work.
OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol (high). Estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used in Astra’s development reached 63.7%. On one example task, Astra met about 94% of criteria. These figures describe performance on the selected research tasks, not general accuracy across Ironclad customers’ contracts.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The reported score should not be read as a completion rate or a measure of safe deployment. It is the average share of rubric criteria met. In a contract workflow, missing a single required approval can matter more than getting several other steps right. OpenAI’s example describes procurement rules that require Finance approval above a spending threshold, Security review for some requests and Legal review for nonstandard terms. A workflow that misses one of those checks could send a purchase forward without a required review.
The results point to both a possible benefit and a constraint. Agents may eventually handle routine work inside specialized software, but businesses need to know which requirements the system misses and how those failures are caught. OpenAI’s own account says human oversight remains relevant when an agent may lose track of a business rule during a multi-step task. In high-consequence settings, an average score alone does not establish that a workflow is dependable.
The collaboration also positions software companies as potential training partners. OpenAI said it is inviting a small number of vendors to bring difficult tasks, domain experts, secure test environments and research-appropriate data. For vendors, stronger agents could make their products more useful. At the same time, if customers increasingly act through agents rather than screens, a product’s lasting value may depend on its underlying rules, records and controls—not just its interface.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
OpenAI’s October 6 publication covered two projects, including a large collection of mathematics manuscripts. The Ironclad post received less attention, but it described a different kind of model development: training and evaluation inside a software vendor’s product and on workflows chosen by people familiar with that work. “Ironclad” in this case refers to the contract-management company, not a new AI agent framework.
The collaboration’s stated aim was to train models to understand business rules, carry out multi-step tasks in specialized software and check finished work against the original requirements. OpenAI said the tasks were built from public contract filings and tested in hosted copies of Ironclad. Its account does not describe the result as a general benchmark of contract work, and the 11 tasks are a limited research set.
The time figures need separate treatment from the rubric results. OpenAI’s footnote says the estimates are based on assumed processing and generation speeds, rather than observed time savings for customers. The estimates apply to the research tasks, not to Ironclad workflows as a whole. The comparison with a human’s estimated 30-to-40-minute task time is therefore not evidence that businesses can currently cut contract processing time by a particular amount.
What the Reported Scores Leave Open
The publication does not establish how the model would perform across the full range of Ironclad customers, contract types or operational conditions. The reported average also does not show, by itself, which individual criteria failed on each task or how often a missed requirement would create a consequential error. One showcase task with about 94% of criteria met does not resolve those questions for the other workflows.
It is also unclear from the reported figures how the model would perform under production conditions, including on changing business rules, unusual contract language or incomplete information. The source material describes the work as research and says human oversight still matters; it does not provide evidence that the model is ready to run these workflows autonomously. No measured customer time savings are reported, and OpenAI’s simulated estimates should not be treated as real-world productivity results.
OpenAI said it used no non-public Ironclad customer data, but the account does not specify every detail of the evaluation protocol or provide a broader external replication. Further results would be needed to judge reliability beyond this small task set.
Further Tests Before Business Use
OpenAI said it is seeking a small number of software-company partners to identify tasks current agents cannot reliably complete. Its requested ingredients include concrete examples of failures, people with detailed knowledge of the work, a secure environment for testing and data that can safely be used for research. The company has not, in the supplied material, announced a broader rollout date or named additional partners.
For businesses considering agents in contract or procurement systems, the next useful evidence would be task-level results: what requirements were missed, how those failures were detected, and how consistently the system handles varied cases. Companies would also need to establish where human review is mandatory and how approvals and audit records are preserved. The current report describes a direction for research, not a demonstrated replacement for experienced staff or established controls.
Key Questions
What did OpenAI and Ironclad test?
They tested models on 11 legal, commercial and procurement workflows in hosted copies of Ironclad’s contract-management product, including nondisclosure agreements and procurement approvals.
Does a 55% score mean the model completed 55% of tasks?
No. OpenAI reported that GPT-6 Astra met an average 55% of rubric criteria across the tasks. The figure is not the percentage of tasks completed and does not show that each workflow is safe to use.
Did the model cut customer work time in half?
The report does not show measured customer savings. OpenAI described the 19.2-minute estimate for Astra as simulated, based on assumed processing and generation speeds, and compared it with estimated times for the research tasks.
What data did OpenAI say it used?
OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can companies use the agent without human review?
The supplied report does not establish that. OpenAI says human oversight remains important, and the reported average score leaves open which business rules the model may miss on a given task.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
