🔍 Read the full analysis: Is Your Software Part Of OpenAI Agent Training? Review Ironclad’s Fine Print on ThorstenMeyerAI.com
Get privacy and security gear delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training a frontier model on 11 contract-management tasks in hosted Ironclad software, using synthetic tasks based on public SEC filings. GPT-6 Astra met an average 55% of task criteria, and OpenAI says its time estimates were simulated rather than measured customer savings. The work also signals OpenAI is seeking software partners to help train agents on difficult professional workflows.
OpenAI said on October 6 that it trained its frontier model, GPT-6 Astra, on contract-management tasks inside Ironclad’s hosted software, reporting an average of 55% of evaluation criteria met across 11 tasks. The post also invites other software companies to work with OpenAI on agent training, while its results and simulated time estimates do not establish that agents are ready to handle contract workflows without human review.
OpenAI and Ironclad selected 11 tasks spanning legal, commercial and procurement work. Examples included preparing nondisclosure agreements, creating procurement approval processes and revising a reusable clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was graded against a rubric of 8 to 50 criteria, depending on complexity; the score represents criteria met, not tasks fully completed.
For training, Ironclad supplied hosted copies of its product where models could practise. OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI identified GPT-6 Astra as the first frontier model trained through this approach.
OpenAI reported an average criteria score of 55.0% for GPT-6 Astra, compared with 41.6% for GPT-5.6 Sol at the high setting. Astra’s estimated time per attempt was 19.2 minutes, against 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of the criteria on one showcase task. The post does not describe that single result as representative of the full set.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The reported score points to both progress and a practical limit. Meeting 55% of rubric criteria is not the same as completing 55% of tasks successfully. A procurement workflow might require Finance approval above a spending threshold, a Security review for specified requests and Legal review for nonstandard terms. Missing one required control could make the resulting process unsafe or unusable, even if other steps were handled correctly.
That distinction matters to organizations considering agents for contracts, procurement and other controlled work. OpenAI’s results are research measurements on selected tasks, not evidence of customer productivity gains or reliable production performance. The company’s own discussion acknowledges the need for human oversight when an agent may lose track of a business rule. Buyers will need to know not only an overall score, but which requirements fail, how errors are caught and who remains accountable for approving the result.
For software vendors, the partnership proposal carries a business question as well as a technical one. Training agents to work inside a product may make it more useful, but agents that can operate its interface could also become the main way customers interact with it. In that scenario, a vendor’s lasting value may depend on its business rules, records, audit trail and controls, not just its screens. That is an implication of the approach, not a stated outcome of the Ironclad research.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
The OpenAI post, titled “Advancing computer use with Ironclad,” describes a collaboration focused on agents performing multi-step work in specialized business software. Tasks were chosen by Ironclad staff and OpenAI employees who use the product, and evaluated against task-specific criteria. This setup is narrower than a broad test of all Ironclad features or all contract-management work.
The reported time figures need particular care. OpenAI’s footnote says the 19.2- and 37-minute estimates are simulated, based on assumed processing and generation speeds. They are not observed time savings for customers, and the 11 research tasks do not stand in for Ironclad workflows generally. Comparing those estimates with the 30-to-40-minute estimate for an experienced user does not establish that an agent completes the work faster or more accurately in practice.
OpenAI also described the Ironclad work as a model for future collaborations. It said it is seeking a small number of software-company partners willing to supply a concrete example of a task current agents struggle with, subject-matter experts, a secure testing environment and data suitable for research. The post frames the work as testing how agents can respect business rules within a full software platform, rather than replacing the platform outright.
As an affiliate, we earn on qualifying purchases.
What the Published Scores Cannot Show
The post does not establish how GPT-6 Astra would perform across a broader range of customer work, on live contracts or under day-to-day operational conditions. It also does not provide enough detail in the supplied material to identify which individual criteria were missed across all 11 tasks, how consistently the results held across repeated attempts, or how the model compared with human users completing the same tasks under measured conditions.
OpenAI says it used no non-public Ironclad customer data, but the material does not set out all terms of the partnership, including any future arrangements for data use, model access or commercial deployment. Nor does it provide measured customer time savings. The reported 55% average is a research rubric result, and the simulated time estimates should not be treated as evidence of realized productivity gains.
As an affiliate, we earn on qualifying purchases.
Questions Before Agents Reach Production
OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. The next relevant evidence would include additional partner projects, clearer reporting on task-level failures and results from testing that reflects real operating conditions. The October 6 post does not give a timeline for those projects or announce a customer deployment schedule.
Companies evaluating agents in their own systems can ask vendors to identify the exact evaluation criteria and disclose which requirements failed. They should also establish how a system handles required approvals, how human reviewers can inspect and correct its work, what records are retained for audit, and what data is used for training or evaluation. Until performance is shown on the workflows and controls a buyer actually relies on, the Ironclad results support continued testing—not unattended contract processing.
contract workflow automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Does the 55% score mean GPT-6 Astra completed 55% of the tasks?
No. OpenAI reported that Astra met an average of 55% of the evaluation criteria across the tasks. That is not a task completion rate or a measure of how many workflows were fully correct.
Did OpenAI use private Ironclad customer contracts?
OpenAI said it used synthetic tasks based on publicly filed contracts from the SEC’s EDGAR database and did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.
Are the reported time reductions measured customer savings?
No. OpenAI described the times as simulated estimates based on assumed processing and generation speeds. They are not measured savings achieved by Ironclad customers.
Can companies use agents to manage contracts without human review?
The reported results do not support that conclusion. Astra met an average of 55% of criteria, and OpenAI’s post says human oversight remains relevant when agents may lose track of business rules. The company has not established that these workflows are ready for unsupervised use.
What does OpenAI want from other software companies?
OpenAI said it is seeking a small number of partners to provide a difficult real-world task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not specify a timeline or commercial terms.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
