Is Your Software Part Of OpenAI Agent Training? Review Ironclad’s Fine Print
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Is Your Software Part Of OpenAI Agent Training? Review Ironclad’s Fine Print on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model on 11 contract-management tasks in hosted Ironclad software, using synthetic tasks based on public SEC filings. GPT-6 Astra met an average 55% of task criteria, and OpenAI says its time estimates were simulated rather than measured customer savings. The work also signals OpenAI is seeking software partners to help train agents on difficult professional workflows.

OpenAI said on October 6 that it trained its frontier model, GPT-6 Astra, on contract-management tasks inside Ironclad’s hosted software, reporting an average of 55% of evaluation criteria met across 11 tasks. The post also invites other software companies to work with OpenAI on agent training, while its results and simulated time estimates do not establish that agents are ready to handle contract workflows without human review.

OpenAI and Ironclad selected 11 tasks spanning legal, commercial and procurement work. Examples included preparing nondisclosure agreements, creating procurement approval processes and revising a reusable clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was graded against a rubric of 8 to 50 criteria, depending on complexity; the score represents criteria met, not tasks fully completed.

For training, Ironclad supplied hosted copies of its product where models could practise. OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI identified GPT-6 Astra as the first frontier model trained through this approach.

OpenAI reported an average criteria score of 55.0% for GPT-6 Astra, compared with 41.6% for GPT-5.6 Sol at the high setting. Astra’s estimated time per attempt was 19.2 minutes, against 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of the criteria on one showcase task. The post does not describe that single result as representative of the full set.

At a glance
reportWhen: Published October 6; current performanc…
The developmentOpenAI’s October 6 post described training GPT-6 Astra on contract-management tasks inside Ironclad and invited other software companies to explore similar partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The reported score points to both progress and a practical limit. Meeting 55% of rubric criteria is not the same as completing 55% of tasks successfully. A procurement workflow might require Finance approval above a spending threshold, a Security review for specified requests and Legal review for nonstandard terms. Missing one required control could make the resulting process unsafe or unusable, even if other steps were handled correctly.

That distinction matters to organizations considering agents for contracts, procurement and other controlled work. OpenAI’s results are research measurements on selected tasks, not evidence of customer productivity gains or reliable production performance. The company’s own discussion acknowledges the need for human oversight when an agent may lose track of a business rule. Buyers will need to know not only an overall score, but which requirements fail, how errors are caught and who remains accountable for approving the result.

For software vendors, the partnership proposal carries a business question as well as a technical one. Training agents to work inside a product may make it more useful, but agents that can operate its interface could also become the main way customers interact with it. In that scenario, a vendor’s lasting value may depend on its business rules, records, audit trail and controls, not just its screens. That is an implication of the approach, not a stated outcome of the Ironclad research.

Amazon

contract management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

The OpenAI post, titled “Advancing computer use with Ironclad,” describes a collaboration focused on agents performing multi-step work in specialized business software. Tasks were chosen by Ironclad staff and OpenAI employees who use the product, and evaluated against task-specific criteria. This setup is narrower than a broad test of all Ironclad features or all contract-management work.

The reported time figures need particular care. OpenAI’s footnote says the 19.2- and 37-minute estimates are simulated, based on assumed processing and generation speeds. They are not observed time savings for customers, and the 11 research tasks do not stand in for Ironclad workflows generally. Comparing those estimates with the 30-to-40-minute estimate for an experienced user does not establish that an agent completes the work faster or more accurately in practice.

OpenAI also described the Ironclad work as a model for future collaborations. It said it is seeking a small number of software-company partners willing to supply a concrete example of a task current agents struggle with, subject-matter experts, a secure testing environment and data suitable for research. The post frames the work as testing how agents can respect business rules within a full software platform, rather than replacing the platform outright.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Published Scores Cannot Show

The post does not establish how GPT-6 Astra would perform across a broader range of customer work, on live contracts or under day-to-day operational conditions. It also does not provide enough detail in the supplied material to identify which individual criteria were missed across all 11 tasks, how consistently the results held across repeated attempts, or how the model compared with human users completing the same tasks under measured conditions.

OpenAI says it used no non-public Ironclad customer data, but the material does not set out all terms of the partnership, including any future arrangements for data use, model access or commercial deployment. Nor does it provide measured customer time savings. The reported 55% average is a research rubric result, and the simulated time estimates should not be treated as evidence of realized productivity gains.

Amazon

AI-powered NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions Before Agents Reach Production

OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. The next relevant evidence would include additional partner projects, clearer reporting on task-level failures and results from testing that reflects real operating conditions. The October 6 post does not give a timeline for those projects or announce a customer deployment schedule.

Companies evaluating agents in their own systems can ask vendors to identify the exact evaluation criteria and disclose which requirements failed. They should also establish how a system handles required approvals, how human reviewers can inspect and correct its work, what records are retained for audit, and what data is used for training or evaluation. Until performance is shown on the workflows and controls a buyer actually relies on, the Ironclad results support continued testing—not unattended contract processing.

Amazon

contract workflow automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the 55% score mean GPT-6 Astra completed 55% of the tasks?

No. OpenAI reported that Astra met an average of 55% of the evaluation criteria across the tasks. That is not a task completion rate or a measure of how many workflows were fully correct.

Did OpenAI use private Ironclad customer contracts?

OpenAI said it used synthetic tasks based on publicly filed contracts from the SEC’s EDGAR database and did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.

Are the reported time reductions measured customer savings?

No. OpenAI described the times as simulated estimates based on assumed processing and generation speeds. They are not measured savings achieved by Ironclad customers.

Can companies use agents to manage contracts without human review?

The reported results do not support that conclusion. Astra met an average of 55% of criteria, and OpenAI’s post says human oversight remains relevant when agents may lose track of business rules. The company has not established that these workflows are ready for unsupervised use.

What does OpenAI want from other software companies?

OpenAI said it is seeking a small number of partners to provide a difficult real-world task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not specify a timeline or commercial terms.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
SenseTime-W’s Profitable Results Highlight AI Industry Momentum

SenseTime-W’s Profitable Results Highlight AI Industry Momentum

SenseTime-W posts interim profit of RMB 607M and 28.2% growth in generative AI revenue, signaling a strategic shift and sector momentum amid competitive pressures.
Search as Code: Perplexity Is Right About the Future — Just Not First to It

Search as Code: Perplexity Is Right About the Future — Just Not First to It

Perplexity introduces Search as Code, enabling AI agents to dynamically assemble search pipelines, claiming significant efficiency gains and accuracy improvements.
A War Room for Your Next Idea: Inside IdeaClyst

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst offers founders a local-first AI-driven war room to validate ideas, reduce costly mistakes, and make strategic decisions with confidence.
The Logic Behind OpenAI’s Price Cuts And Unchanged Benchmarks For GPT‑6 Sol And Luna

The Logic Behind OpenAI’s Price Cuts And Unchanged Benchmarks For GPT‑6 Sol And Luna

OpenAI reduces GPT-6 Sol and Luna prices by 50%, maintaining performance benchmarks. Analysis shows cost savings without loss of quality, impacting AI deployment strategies.