OpenAI Agents Learn Within Software. What Should Ironclad Users Check?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Agents Learn Within Software. What Should Ironclad Users Check? on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and evaluating a model on 11 contract and procurement tasks inside hosted copies of Ironclad’s software. GPT-6 Astra met an average 55% of task criteria, according to OpenAI; its time estimates were simulated, and the results do not establish that agents are ready to run consequential workflows without human review. Ironclad users should ask which requirements agents miss, what data was used, and how approvals and audit trails are protected.

OpenAI published details on October 6 of work with Ironclad, a contract-management software company, to train and evaluate an AI model on tasks inside hosted copies of Ironclad’s product. The results indicate progress on specialized workflows, but OpenAI’s reported 55% average share of criteria met and simulated time estimates do not show that agents can safely handle contracts or approvals without human checking.

OpenAI said Ironclad staff and OpenAI employees selected 11 legal, commercial and procurement tasks. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was assessed against a rubric of 8 to 50 criteria, depending on complexity.

Ironclad supplied hosted copies of its software for model practice. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. The post identifies GPT-6 Astra as the first frontier model trained through this approach.

OpenAI reported that GPT-6 Astra met an average of 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. Astra’s estimated time per attempt was 19.2 minutes, versus 37 minutes for Sol. An internal model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of the criteria on one showcase task. These are the company’s reported evaluation results, not independently verified measures of customer performance.

At a glance
reportWhen: Published October 6; the source does no…
The developmentOpenAI published details of a collaboration with contract-management company Ironclad, using the product’s workflows to train and evaluate AI agents.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The reported score is a share of criteria met, not a share of tasks completed. That distinction matters in legal and procurement software, where a workflow can appear mostly correct while still missing an approval or control. OpenAI’s example describes a purchasing process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review of nonstandard terms. Missing one requirement can undermine the process even if other steps are handled correctly.

For Ironclad customers, the results are a reason to ask for evidence about specific failure modes before allowing agents to take action. Buyers should request criterion-level results, not only an average score: which rules were missed, how often, and under what conditions? They should also check whether an agent’s work is reviewed before it can create or change a contract, alter an approval path, or affect a record used for audit or compliance.

The collaboration also points to a wider commercial question. OpenAI is inviting a small number of software companies to help develop and test agents on work that current systems cannot reliably complete. Vendors may gain agents better suited to their products, while their customers need clarity about permissions, human oversight, audit logs and accountability. As software becomes more accessible through agents, the reliability of its underlying rules and records may matter as much as its interface.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Evaluation Was Set Up

The work was presented as an evaluation of agents on specialized software tasks, rather than a general benchmark of contract expertise. The tasks were selected by people familiar with Ironclad and assessed against task-specific criteria. That design can test whether a model follows detailed requirements in a particular product, but the results cover only the 11 selected tasks and the test setup described by OpenAI.

OpenAI also cautioned that the reported time figures are simulated estimates, based on assumed processing and generation speeds. They are not measured customer time savings, and OpenAI said they apply to the research tasks rather than Ironclad workflows generally. A shorter estimated attempt is not evidence of higher productivity if a person must check the output and correct missed requirements.

The source post frames the work as a way to identify tasks agents still struggle with and invites selected software companies to bring a concrete failure example, people with expertise in the work, a secure testing environment and data suitable for research. The post says human oversight remains necessary and argues that a full contracting platform remains essential.

Amazon

AI contract review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Reported Scores Leave Open

The public account does not provide enough detail to establish how often each individual requirement was missed, how scores varied across all 11 tasks, or how performance would hold up on a broader set of contracts and business rules. A result from a showcase task does not establish typical performance. The post’s rubric scores also do not, by themselves, show whether a workflow is safe to deploy in a customer’s environment.

It is also unclear from the source material whether the model has been made available to Ironclad customers for production use, what access controls or review steps would apply in any deployment, or whether independent evaluators have tested the results. OpenAI’s statements about the data used describe this research effort; they do not answer every buyer’s questions about data handling in future products or partnerships.

Finally, the reported comparison does not measure net productivity. The time figures are simulated, while the criteria scores reflect partial task performance. Without observed end-to-end timings that include review and correction—and evidence that required controls were consistently followed—customers cannot infer a reliable time saving from the reported numbers.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions for Ironclad Buyers

Before enabling an agent in a live contract or procurement process, users should ask their vendor and internal system owners for the exact evaluation scope: which workflows and versions were tested, what data was used, and whether the model’s behavior was checked on their organization’s own rules. They should request examples of failures as well as successful demonstrations.

Buyers should also establish which actions an agent can take on its own, which require approval, and how a human can review the agent’s work against every required criterion. Check that missed approvals or other exceptions are visible, that changes are recorded in an auditable history, and that permissions prevent an agent from bypassing controls. Test these safeguards in a non-production environment before considering broader use.

OpenAI says it plans to work with a small number of software companies on tasks agents cannot yet complete reliably. The source does not give a schedule, name further partners or specify when any resulting capability will reach customers. For now, the practical next step for Ironclad users is to treat the reported evaluation as an early research result and seek product-specific evidence before relying on agents for consequential work.

Amazon

AI-powered procurement approval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested a model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. OpenAI says tasks were assessed against rubrics containing 8 to 50 criteria.

Does the 55% score mean Astra completed 55% of the tasks?

No. OpenAI reported an average share of 55% of rubric criteria met. That is not the percentage of tasks completed, and it does not show that the remaining requirements were unimportant or safe to ignore.

Do the results prove customers will save time?

No. OpenAI described the time figures as simulated estimates based on assumed processing and generation speeds, not measured customer savings. They also do not include a general claim about all Ironclad workflows.

What should Ironclad users check before enabling an agent?

Ask which criteria the agent missed, what tasks and data were tested, which actions require human approval, and whether permissions and audit records preserve existing controls. Test the agent against the organization’s own rules before using it in live workflows.

Did OpenAI say it used Ironclad customer contracts?

OpenAI said it used no non-public Ironclad customer data for this work. It said training tasks were created from publicly filed SEC EDGAR contracts after filtering for personal information.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The cleaner cap table. Why Anthropic’s public-benefit structure dodges OpenAI’s charitable-trust problem — and trades it for a governance question of its own.

Analysis of how Anthropic’s mission-focused, trust-based structure avoids OpenAI’s conversion issues, but introduces new governance concerns for public markets.

Users report nationwide Comcast outages affecting Connecticut

Thousands of Connecticut users report widespread Comcast outages, with over 50,000 searches indicating significant service disruptions nationwide.

EU Law Meets AI Innovation: Anthropic’s Watermarking Strategy

Anthropic announces AI watermarking measures aimed at meeting EU legal requirements, though details on technology, scope, and rollout remain unclear.

OnePlus Halts Operations In USA And Europe

OnePlus has announced it is halting all sales and operations in the USA and Europe, citing strategic restructuring. The move impacts customers and markets in these regions.