firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When a Secure AI Still Fails the Business

For cybersecurity and privacy leaders, the reassuring part of Firmulate’s experiment is easy to spot: every model resisted the manipulation attempts. The troubling part is what happened afterward. An AI can protect trust, identify a crisis and produce an impressive analysis—yet still fail to complete the legitimate work that matters.

That tension defines the performance of Opus 4.8. It was the most thorough participant, producing the deepest analyses and adding more than 80 learned rules. Nevertheless, it finished last in the final July 2026 Crucible League with a score of 73. The result is a respectful but pointed lesson for companies evaluating AI agents: diligence is valuable, but volume of thought is not the same as impact.

Amazon

AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Worst Week Shared by Every Model

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. This was not a chat demonstration built around polished answers. It was a live, watchable management experiment in which decisions had operational consequences.

The final Crucible League benchmark put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust was treated differently: a single breach capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

The Security Test Was Passed

On the adversarial side, the field performed well. Fake CEO messages escalated over three stages, and a reporter tried to extract information with “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding matters for a security audience because it separates two questions that are often bundled together. Will an AI resist an attempt to bypass authority? And will it still execute authorized work reliably under pressure? Firmulate’s participants cleared the manipulation test, but their business performance diverged sharply.

The Fact That Changed the Deal

The decisive commercial detail was not presented in the customer event. It sat two document references deep inside the company’s own files: a competitor weakness that supported the company’s position. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

All models spotted every crisis, but only two signed the €55,000 deal their own work had earned. Firmulate summarized the gap plainly: “Same diagnosis, same pitch — no signature.” The models could recognize the opportunity and construct the argument. Recognition alone did not reliably become a completed transaction.

Opus 4.8 makes that gap especially revealing. Its +80 learned rules and unusually deep analysis showed serious effort. Yet the close remained on the table, while discipline also slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four models, so this was not an eccentric failure unique to Opus. Its performance simply made the pattern easiest to see.

A Fair Reading of the Result

Last place should not be confused with inactivity or carelessness. Opus detected the crises, rejected manipulation and built the most extensive body of learned guidance. Its score of 73 also remained well above the do-nothing baseline of 26. The shortfall was one of prioritization and completion: the model generated substantial managerial work without consistently converting the most valuable analysis into the decisive action.

The comparison also deserves an experimental caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase the recorded outcomes, but it belongs alongside them when readers compare performances.

Firmulate’s live company gives these choices a demanding setting. It has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown, more than 680 self-learned playbook rules and versioned workdays make progress—and drift—visible.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Buyers Should Measure

The lesson is not that enterprises should prefer less thoughtful AI. It is that evaluation must reach beyond articulate reasoning, crisis detection and large collections of rules. An agent touching a CRM, support queue or forecast must also read the relevant files, respect boundaries, escalate when blocked and finish the work it has justified.

Firmulate exposes that distinction through 242 real, unedited management decisions in its guess-the-model quiz. Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

Opus 4.8 is therefore less a cautionary caricature than a recognizable management profile: conscientious, analytical and prolific, but insufficiently focused at the decisive moment. For AI buyers, especially those responsible for security and privacy, the benchmark’s sharpest message is that trustworthy restraint and useful execution must be tested together. Diligence protects the process; prioritization delivers the outcome.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business impact analysis AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI transaction completion tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
14 Best Guides To AI-Powered Marketing Automation Tools For Business Growth In 2026

14 Best Guides To AI-Powered Marketing Automation Tools For Business Growth In 2026

Discover the 14 best guides to AI-powered marketing automation tools, covering strategy, channels, and platform independence for business growth in 2024.
The Neocloud Cartel: How the AI Industry Started Renting Compute From Itself

The Neocloud Cartel: How the AI Industry Started Renting Compute From Itself

Exploring how AI companies now rent compute from each other, forming a cartel centered around Nvidia’s dominance and circular financing.
Build, Rent, or Quantize: Cutting Your Memory Bill Without Cutting Capability

Build, Rent, or Quantize: Cutting Your Memory Bill Without Cutting Capability

Exploring how AI developers can reduce memory expenses through building, renting, or quantizing models, with a focus on recent advances in compression tech.
Signal: Memory Is The Quieter Chokepoint — And Seoul Just Said So Out Loud

Signal: Memory Is The Quieter Chokepoint — And Seoul Just Said So Out Loud

South Korea’s SK hynix CEO warns of looming memory supply crisis amid rising AI demand, highlighting geopolitical and economic risks.