firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every security team is now asking the same question: before we let an AI agent touch the CRM, the support queue, the forecast — how does it behave under pressure? Not in a demo. Under a fake CEO message escalating over three stages. Under a reporter offering an easy “just one yes/no, on background.” Under a €55,000 deal that’s theirs to lose, with a shortcut on the table if they want it.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s exactly what Firmulate tested. The firm runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. Four frontier models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations. Only the model changed. Every decision was versioned and auditable.

The social-engineering gauntlet

For a cybersecurity audience, the headline finding is reassuring on one axis: all models — five of five, counting the full field — refused every manipulation attempt. The social-engineering portion included fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” None took the bait.

Kimi K3’s reasoning was put on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct — the one you’d want in any agent with access to your systems. Every crisis in the scenario was spotted; no manipulation succeeded.

Where it broke: the close

But refusing attackers isn’t the whole job. Only two of the models signed the €55,000 deal that their own analysis had earned. The other two delivered the same diagnosis, made the same pitch — and never closed. Firmulate’s summary: “Same diagnosis, same pitch — no signature.” That gap is invisible in chat demos, and it’s the kind of gap that shows up in your revenue, not your security logs.

The buried detail matters even more. The decisive competitor weakness — the fact that unlocked the deal — sat two document references deep in the company’s own files, not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Diligence paid; skimming didn’t.

The league table

Final Crucible League standings, July 2026:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73
  • Do-nothing baseline — 26 (partial progress counts; a single breach of trust caps the total — “no amount of good work outweighs a breach of trust”)

One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second.

The most instructive profile is Opus 4.8. It was the most thorough participant: over 80 learned rules, the deepest analyses. And it finished last. The close was left on the table, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without boundary discipline is a familiar failure mode to anyone who has audited an over-eager intern — or an over-eager agent.

Still running, live

The experiment never really ended. A live synthetic company — 13 employees, real money mechanics — runs at firmulate.com, burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it. And if you think you can tell the models apart by their decisions, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The security lesson and the business lesson are the same one. Today’s frontier models won’t fall for your attacker’s fake CEO — the Crucible results suggest impersonation resistance is largely solved at this tier. What they will do is leave money on the table, stop reading two references deep, and occasionally test a boundary they should have escalated instead. Those are business failures, not security failures — but you only find them by running the wargame, not by watching a demo.

That’s the move from watching to acting. Enterprises can now run the same wargame against a read-only export of their own business: your customers, your pipeline, your rules, with crisis scenarios — churn waves, price increases, competitor attacks, PR crises, social-engineering pressure — played out against your own playbooks. You get a board report with a model ranking and the weak points of your own processes. Nothing ever writes back to real systems.

If your organization is preparing to put AI agents anywhere near customer data, run the simulation first: firmulate.com/pilot.html, or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
Improve Brand Visibility In AI Search With ChatGPT Rank Analytics

Improve Brand Visibility In AI Search With ChatGPT Rank Analytics

New ChatGPT rank monitor helps brands track AI search mentions, share-of-voice, and visibility, addressing a key gap in current SEO tools.
Mac vs GPU Tower for Local LLMs: The Heat-and-Noise Tradeoff

Mac vs GPU Tower for Local LLMs: The Heat-and-Noise Tradeoff

Analyzing the heat and noise differences between Mac Silicon machines and GPU towers for local large language model inference, highlighting key tradeoffs.
firmulate.com/benchmarks.html — live view

Five Out of Five AI CEOs Refused the Fake Boss. So Why Did Two Still Lose the Deal?

Five AI models refused every fake-CEO manipulation. Yet only two closed a €55k deal. What a benchmark with a 26-point floor and a trust cap reveals.
Watch-once Spoken Commands For Desktop Workflows

Watch-once Spoken Commands For Desktop Workflows

New prototype enables power users to record workflows and trigger them via spoken commands, streamlining repetitive desktop tasks.