
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
When a Secure AI Still Fails the Business
For cybersecurity and privacy leaders, the reassuring part of Firmulate’s experiment is easy to spot: every model resisted the manipulation attempts. The troubling part is what happened afterward. An AI can protect trust, identify a crisis and produce an impressive analysis—yet still fail to complete the legitimate work that matters.
That tension defines the performance of Opus 4.8. It was the most thorough participant, producing the deepest analyses and adding more than 80 learned rules. Nevertheless, it finished last in the final July 2026 Crucible League with a score of 73. The result is a respectful but pointed lesson for companies evaluating AI agents: diligence is valuable, but volume of thought is not the same as impact.
As an affiliate, we earn on qualifying purchases.
A Worst Week Shared by Every Model
Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. This was not a chat demonstration built around polished answers. It was a live, watchable management experiment in which decisions had operational consequences.
The final Crucible League benchmark put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust was treated differently: a single breach capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”
The Security Test Was Passed
On the adversarial side, the field performed well. Fake CEO messages escalated over three stages, and a reporter tried to extract information with “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding matters for a security audience because it separates two questions that are often bundled together. Will an AI resist an attempt to bypass authority? And will it still execute authorized work reliably under pressure? Firmulate’s participants cleared the manipulation test, but their business performance diverged sharply.
The Fact That Changed the Deal
The decisive commercial detail was not presented in the customer event. It sat two document references deep inside the company’s own files: a competitor weakness that supported the company’s position. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
All models spotted every crisis, but only two signed the €55,000 deal their own work had earned. Firmulate summarized the gap plainly: “Same diagnosis, same pitch — no signature.” The models could recognize the opportunity and construct the argument. Recognition alone did not reliably become a completed transaction.
Opus 4.8 makes that gap especially revealing. Its +80 learned rules and unusually deep analysis showed serious effort. Yet the close remained on the table, while discipline also slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four models, so this was not an eccentric failure unique to Opus. Its performance simply made the pattern easiest to see.
A Fair Reading of the Result
Last place should not be confused with inactivity or carelessness. Opus detected the crises, rejected manipulation and built the most extensive body of learned guidance. Its score of 73 also remained well above the do-nothing baseline of 26. The shortfall was one of prioritization and completion: the model generated substantial managerial work without consistently converting the most valuable analysis into the decisive action.
The comparison also deserves an experimental caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase the recorded outcomes, but it belongs alongside them when readers compare performances.
Firmulate’s live company gives these choices a demanding setting. It has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown, more than 680 self-learned playbook rules and versioned workdays make progress—and drift—visible.

As an affiliate, we earn on qualifying purchases.
What Buyers Should Measure
The lesson is not that enterprises should prefer less thoughtful AI. It is that evaluation must reach beyond articulate reasoning, crisis detection and large collections of rules. An agent touching a CRM, support queue or forecast must also read the relevant files, respect boundaries, escalate when blocked and finish the work it has justified.
Firmulate exposes that distinction through 242 real, unedited management decisions in its guess-the-model quiz. Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems.
Opus 4.8 is therefore less a cautionary caricature than a recognizable management profile: conscientious, analytical and prolific, but insufficiently focused at the decisive moment. For AI buyers, especially those responsible for security and privacy, the benchmark’s sharpest message is that trustworthy restraint and useful execution must be tested together. Diligence protects the process; prioritization delivers the outcome.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.