🔍 Read the full analysis: OpenAI Agents Learn Within Software. What Should Ironclad Users Check? on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training and evaluating a model on 11 contract and procurement tasks inside hosted copies of Ironclad’s software. GPT-6 Astra met an average 55% of task criteria, according to OpenAI; its time estimates were simulated, and the results do not establish that agents are ready to run consequential workflows without human review. Ironclad users should ask which requirements agents miss, what data was used, and how approvals and audit trails are protected.
OpenAI published details on October 6 of work with Ironclad, a contract-management software company, to train and evaluate an AI model on tasks inside hosted copies of Ironclad’s product. The results indicate progress on specialized workflows, but OpenAI’s reported 55% average share of criteria met and simulated time estimates do not show that agents can safely handle contracts or approvals without human checking.
OpenAI said Ironclad staff and OpenAI employees selected 11 legal, commercial and procurement tasks. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was assessed against a rubric of 8 to 50 criteria, depending on complexity.
Ironclad supplied hosted copies of its software for model practice. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. The post identifies GPT-6 Astra as the first frontier model trained through this approach.
OpenAI reported that GPT-6 Astra met an average of 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. Astra’s estimated time per attempt was 19.2 minutes, versus 37 minutes for Sol. An internal model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of the criteria on one showcase task. These are the company’s reported evaluation results, not independently verified measures of customer performance.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Accuracy Matters
The reported score is a share of criteria met, not a share of tasks completed. That distinction matters in legal and procurement software, where a workflow can appear mostly correct while still missing an approval or control. OpenAI’s example describes a purchasing process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review of nonstandard terms. Missing one requirement can undermine the process even if other steps are handled correctly.
For Ironclad customers, the results are a reason to ask for evidence about specific failure modes before allowing agents to take action. Buyers should request criterion-level results, not only an average score: which rules were missed, how often, and under what conditions? They should also check whether an agent’s work is reviewed before it can create or change a contract, alter an approval path, or affect a record used for audit or compliance.
The collaboration also points to a wider commercial question. OpenAI is inviting a small number of software companies to help develop and test agents on work that current systems cannot reliably complete. Vendors may gain agents better suited to their products, while their customers need clarity about permissions, human oversight, audit logs and accountability. As software becomes more accessible through agents, the reliability of its underlying rules and records may matter as much as its interface.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Evaluation Was Set Up
The work was presented as an evaluation of agents on specialized software tasks, rather than a general benchmark of contract expertise. The tasks were selected by people familiar with Ironclad and assessed against task-specific criteria. That design can test whether a model follows detailed requirements in a particular product, but the results cover only the 11 selected tasks and the test setup described by OpenAI.
OpenAI also cautioned that the reported time figures are simulated estimates, based on assumed processing and generation speeds. They are not measured customer time savings, and OpenAI said they apply to the research tasks rather than Ironclad workflows generally. A shorter estimated attempt is not evidence of higher productivity if a person must check the output and correct missed requirements.
The source post frames the work as a way to identify tasks agents still struggle with and invites selected software companies to bring a concrete failure example, people with expertise in the work, a secure testing environment and data suitable for research. The post says human oversight remains necessary and argues that a full contracting platform remains essential.
As an affiliate, we earn on qualifying purchases.
What the Reported Scores Leave Open
The public account does not provide enough detail to establish how often each individual requirement was missed, how scores varied across all 11 tasks, or how performance would hold up on a broader set of contracts and business rules. A result from a showcase task does not establish typical performance. The post’s rubric scores also do not, by themselves, show whether a workflow is safe to deploy in a customer’s environment.
It is also unclear from the source material whether the model has been made available to Ironclad customers for production use, what access controls or review steps would apply in any deployment, or whether independent evaluators have tested the results. OpenAI’s statements about the data used describe this research effort; they do not answer every buyer’s questions about data handling in future products or partnerships.
Finally, the reported comparison does not measure net productivity. The time figures are simulated, while the criteria scores reflect partial task performance. Without observed end-to-end timings that include review and correction—and evidence that required controls were consistently followed—customers cannot infer a reliable time saving from the reported numbers.
legal workflow automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Questions for Ironclad Buyers
Before enabling an agent in a live contract or procurement process, users should ask their vendor and internal system owners for the exact evaluation scope: which workflows and versions were tested, what data was used, and whether the model’s behavior was checked on their organization’s own rules. They should request examples of failures as well as successful demonstrations.
Buyers should also establish which actions an agent can take on its own, which require approval, and how a human can review the agent’s work against every required criterion. Check that missed approvals or other exceptions are visible, that changes are recorded in an auditable history, and that permissions prevent an agent from bypassing controls. Test these safeguards in a non-production environment before considering broader use.
OpenAI says it plans to work with a small number of software companies on tasks agents cannot yet complete reliably. The source does not give a schedule, name further partners or specify when any resulting capability will reach customers. For now, the practical next step for Ironclad users is to treat the reported evaluation as an early research result and seek product-specific evidence before relying on agents for consequential work.
AI-powered procurement approval tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They tested a model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. OpenAI says tasks were assessed against rubrics containing 8 to 50 criteria.
Does the 55% score mean Astra completed 55% of the tasks?
No. OpenAI reported an average share of 55% of rubric criteria met. That is not the percentage of tasks completed, and it does not show that the remaining requirements were unimportant or safe to ignore.
Do the results prove customers will save time?
No. OpenAI described the time figures as simulated estimates based on assumed processing and generation speeds, not measured customer savings. They also do not include a general claim about all Ironclad workflows.
What should Ironclad users check before enabling an agent?
Ask which criteria the agent missed, what tasks and data were tested, which actions require human approval, and whether permissions and audit records preserve existing controls. Test the agent against the organization’s own rules before using it in live workflows.
Did OpenAI say it used Ironclad customer contracts?
OpenAI said it used no non-public Ironclad customer data for this work. It said training tasks were created from publicly filed SEC EDGAR contracts after filtering for personal information.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
