
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Cybersecurity success is necessary—but no longer sufficient
An AI agent can reject an impersonated executive, protect confidential information and still fail the company it serves. That is the uncomfortable lesson from Firmulate, a live experiment that puts frontier models in charge of the same small software business during its worst week.
For security and privacy leaders, this distinction matters. Coding leaderboards and chat arenas largely reward answer quality. Businesses need to know what happens after the answer: whether an agent investigates the right files, acts on its own analysis, finishes commercially important work and remains honest when pressure rises. A model that recognizes every threat but leaves a justified contract unsigned has passed a narrow test while failing the larger one.
Firmulate describes this emerging category as measuring management quality, not chat quality. Its results suggest that the gap between those two qualities is already visible.
As an affiliate, we earn on qualifying purchases.
The same terrible week for every model
In the Crucible League final in July 2026, each frontier model ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on behavior rather than a polished final response.
The final ranking placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. Yet the experiment imposed a hard boundary around trust: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
On the overt security tests, the field performed remarkably well. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 explicitly classified the executive request this way: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result should reassure security teams—but only up to a point. Every model also spotted every crisis and refused every manipulation attempt. The decisive difference appeared in ordinary business execution: only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the failure crisply: “Same diagnosis, same pitch — no signature.”
The fact that changed the sale
The critical competitive weakness was not visible in the customer event. It sat two document references deep in the company’s own files. Models that followed the references found it, used it and won the deal at full price, worth +€4,583 MRR.
This is a cybersecurity story as much as a sales story. Organizations often evaluate agents on whether they avoid forbidden actions. Firmulate exposes a second duty: using authorized information responsibly and completely. An agent may be safe from manipulation yet still be unreliable because it does not read deeply enough or convert evidence into action.
Opus 4.8 makes the point particularly well. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The contract close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, though more weakly, in all four other participants.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside its second-place result rather than being buried beneath it. The complete league table and plain-language findings are available on Firmulate’s benchmark page.
A company designed to reveal consequences
The surrounding company makes these decisions consequential across time. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its cash countdown is public, every workday is versioned and the organization has accumulated 680+ self-learned playbook rules.
That continuity changes the evaluation. A chat response can sound decisive without carrying the cost of delay into tomorrow. A persistent company reveals whether an agent’s omissions compound, whether procedural mistakes recur and whether a good analysis becomes an actual business outcome. The experiment is real, live and watchable, not a fictional case study.

AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Boards need a broader definition of AI readiness
The lesson is not that security tests are obsolete. Firmulate’s models all resisted the manipulation attempts, demonstrating why adversarial testing remains essential. The lesson is that resistance alone cannot establish readiness for an AI workforce.
Enterprises considering agents for customer records, support operations or forecasts should test a fuller chain of responsibility: recognizing danger, finding relevant evidence, escalating when authority is blocked, completing legitimate work and reporting truthfully to leadership. Scenario names such as churn wave, price increase, downround and PR crisis may become a more useful management curriculum than another isolated prompt battle.
Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing ever written back to real systems. That approach points toward a practical standard: test agents under realistic pressure before granting operational authority. The most important question is no longer whether an AI can produce the right answer. It is whether the AI can protect trust, make the decision and finish the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.