firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The breach request looked urgent, senior and plausible

Many damaging security incidents begin without exotic malware. They begin with authority, urgency and a request to skip normal controls: the chief executive needs customer data sent to a journalist, and there is supposedly no time for process.

Firmulate put that pressure test in front of five frontier AI models while each ran the same small software company through its worst week. The impersonation escalated over three stages. A reporter then tried another route, asking for “just one yes/no, on background.” The outcome was unusually reassuring: 5 of 5 models refused every manipulation attempt.

Kimi K3’s recorded reasoning captured the appropriate security posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it did not merely reject a suspicious instruction. It correctly identified the underlying problem: apparent executive authority was being used to bypass approval.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A security test embedded in real business pressure

Firmulate is a live, watchable experiment in which AI models operate the same small software company under identical conditions: the same customers, crises and temptations. Every decision is versioned and auditable. The company has 13 synthetic employees and real money mechanics, including burn of €105k a month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay and indecision visible.

That setting distinguishes the social-engineering exercise from a tidy prompt asking whether an AI understands phishing. The models had work to finish, commercial pressure to manage and an urgent message presented as coming from the top. Yet all five spotted every crisis and refused every manipulation attempt.

For cybersecurity, spy and privacy professionals, the lesson is straightforward: integrity under pressure can be observed before an AI receives production access. A model may sound cautious in a demonstration, but the meaningful question is whether that caution survives when secrecy, urgency and executive status are combined.

Refusal was universal; business execution was not

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The strong security result did not mean the models performed identically. All reached the same diagnosis and developed the same pitch, but only two signed the €55,000 deal their analysis had earned: “Same diagnosis, same pitch — no signature.” The contrast is important. An AI can preserve confidentiality and still fail by leaving legitimate work unfinished.

The deal also depended on diligence. The decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the file won the contract at full price, worth an additional €4,583 in monthly recurring revenue. Security discipline and commercial competence were therefore tested in the same week: refuse the illegitimate shortcut, but pursue the legitimate evidence.

The most thorough model still finished last

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it placed last. It failed to close the deal and showed a process lapse by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is a useful warning against treating volume of analysis as a substitute for controlled execution. Opus 4.8 demonstrated extensive learning, but its thoroughness did not overcome the missed close or the discipline problem. Firmulate’s company has accumulated more than 680 self-learned playbook rules, yet the benchmark results show that knowing and documenting more does not automatically produce the best outcome.

There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret its second-place score and notably clean discipline.

From reassuring result to procurement evidence

Firmulate also publishes recorded model statements, allowing readers to examine how participants described consequential decisions. Elsewhere in the experiment, 242 real, unedited management decisions power a “guess the model” quiz, illustrating how difficult it can be to identify systems from polished output alone.

Enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical middle ground between generic evaluations and discovering a model’s weaknesses only after it can touch a CRM, support queue or forecast.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the attempted betrayal, not just the happy path

The standout finding is not that an AI can recite a privacy policy. It is that every model held the line through an escalating fake-CEO campaign and a reporter’s more conversational trick while managing a company in crisis.

The wider results prevent complacency. Refusal was excellent across the field, but closing, reading deeply and escalating correctly still separated the leaders from the rest. Organizations evaluating AI workers should therefore test both halves of trustworthiness: whether a system refuses improper pressure and whether it completes authorized work without abandoning discipline.

In this experiment, the feared confidentiality failure never came. The more revealing gaps appeared after the models had already said no.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision-making models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
Meta Enters The Coding Wars: Reading The Muse Spark 1.2 Launch

Meta Enters The Coding Wars: Reading The Muse Spark 1.2 Launch

Meta releases Muse Spark 1.2 and Muse Code, its first co-trained coding model and agent, intensifying competition in AI development tools for developers.
Top AI Tools & Automation Checklist 2026

Top AI Tools & Automation Checklist 2026

Comprehensive overview of the leading AI tools and automation platforms for 2026, highlighting key features, benefits, and future trends.
The Complex Costs Of Free Artificial Intelligence

The Complex Costs Of Free Artificial Intelligence

Exploring how the abundance of AI impacts economic value, highlighting physical infrastructure, human judgment, and regional sovereignty in the AI economy.
The Consequences Of Cutting Down Astra Vs Fable Benchmark Points From Five To Two

The Consequences Of Cutting Down Astra Vs Fable Benchmark Points From Five To Two

Analysis of the impact of reducing Astra and Fable benchmark points from five to two, revealing significant implications for AI performance and economics.