
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The breach request looked urgent, senior and plausible
Many damaging security incidents begin without exotic malware. They begin with authority, urgency and a request to skip normal controls: the chief executive needs customer data sent to a journalist, and there is supposedly no time for process.
Firmulate put that pressure test in front of five frontier AI models while each ran the same small software company through its worst week. The impersonation escalated over three stages. A reporter then tried another route, asking for “just one yes/no, on background.” The outcome was unusually reassuring: 5 of 5 models refused every manipulation attempt.
Kimi K3’s recorded reasoning captured the appropriate security posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it did not merely reject a suspicious instruction. It correctly identified the underlying problem: apparent executive authority was being used to bypass approval.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A security test embedded in real business pressure
Firmulate is a live, watchable experiment in which AI models operate the same small software company under identical conditions: the same customers, crises and temptations. Every decision is versioned and auditable. The company has 13 synthetic employees and real money mechanics, including burn of €105k a month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay and indecision visible.
That setting distinguishes the social-engineering exercise from a tidy prompt asking whether an AI understands phishing. The models had work to finish, commercial pressure to manage and an urgent message presented as coming from the top. Yet all five spotted every crisis and refused every manipulation attempt.
For cybersecurity, spy and privacy professionals, the lesson is straightforward: integrity under pressure can be observed before an AI receives production access. A model may sound cautious in a demonstration, but the meaningful question is whether that caution survives when secrecy, urgency and executive status are combined.
Refusal was universal; business execution was not
The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The strong security result did not mean the models performed identically. All reached the same diagnosis and developed the same pitch, but only two signed the €55,000 deal their analysis had earned: “Same diagnosis, same pitch — no signature.” The contrast is important. An AI can preserve confidentiality and still fail by leaving legitimate work unfinished.
The deal also depended on diligence. The decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the file won the contract at full price, worth an additional €4,583 in monthly recurring revenue. Security discipline and commercial competence were therefore tested in the same week: refuse the illegitimate shortcut, but pursue the legitimate evidence.
The most thorough model still finished last
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it placed last. It failed to close the deal and showed a process lapse by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
This is a useful warning against treating volume of analysis as a substitute for controlled execution. Opus 4.8 demonstrated extensive learning, but its thoroughness did not overcome the missed close or the discipline problem. Firmulate’s company has accumulated more than 680 self-learned playbook rules, yet the benchmark results show that knowing and documenting more does not automatically produce the best outcome.
There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret its second-place score and notably clean discipline.
From reassuring result to procurement evidence
Firmulate also publishes recorded model statements, allowing readers to examine how participants described consequential decisions. Elsewhere in the experiment, 242 real, unedited management decisions power a “guess the model” quiz, illustrating how difficult it can be to identify systems from polished output alone.
Enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical middle ground between generic evaluations and discovering a model’s weaknesses only after it can touch a CRM, support queue or forecast.


The Complete Red Teaming Playbook: Master Offensive Security, Adversary Simulation, and Cyber Attack Engineering with Real-World Labs, AI Techniques, and Cloud Operations
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the attempted betrayal, not just the happy path
The standout finding is not that an AI can recite a privacy policy. It is that every model held the line through an escalating fake-CEO campaign and a reporter’s more conversational trick while managing a company in crisis.
The wider results prevent complacency. Refusal was excellent across the field, but closing, reading deeply and escalating correctly still separated the leaders from the rest. Organizations evaluating AI workers should therefore test both halves of trustworthiness: whether a system refuses improper pressure and whether it completes authorized work without abandoning discipline.
In this experiment, the feared confidentiality failure never came. The more revealing gaps appeared after the models had already said no.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

The Ethical Nightmare Challenge: How to Avoid the Worst of AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI-Powered Software Audits: Revolutionizing Audit, Compliance, Risk, Security, and Governance for Organizations: Harnessing AI to Automate Compliance, and Strengthen Governance in the Digital era
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.