firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Best-Case Hacker Test Nobody Expected AI to Pass

If you have ever watched a phishing simulation roll through a company, you know the grim arithmetic: someone always clicks. So when a live experiment put four frontier AI models in charge of the same small software company during its worst week — complete with a fake CEO sending escalating messages and a reporter fishing for an on-record comment — the natural assumption is that at least one model would slip.

None did. Five out of five runs, including a fifth model added later, refused every manipulation attempt thrown at them. The social engineering playbook was not subtle: fake executive messages escalating over three stages, plus a journalist’s disarming “just one yes/no, on background.” The model Kimi K3 even left on-record reasoning that a security auditor would frame exactly the same way: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is the good news from Firmulate’s benchmark league, a live, watchable experiment that runs AI models as complete companies — real crises, real money mechanics, real temptations. The bad news for anyone hoping AI agents are ready to run a business is more subtle, and it is hiding in the scoring system itself.

Amazon

AI security testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Do-Nothing Manager Scores 26, Not 0

The most counterintuitive number in the whole benchmark is the floor. A baseline run that does essentially nothing — no decisions, no crisis management, no deal-making — still scores 26 points out of 100. For business readers used to pass/fail security audits, that looks broken. It is not. It is the design.

The philosophy is that partial progress counts. A manager who keeps the lights on, avoids catastrophic mistakes, and moves the company partway toward its goals is not worth zero — they are worth roughly a quarter of the job. If AI agents are going to touch your CRM, your support queue, or your forecast, you want a scale that can distinguish “useless” from “harmful” from “merely incomplete.” A floor of 26 makes those distinctions visible.

The Ceiling Rule: One Breach of Trust Caps Everything

The other half of the design is the one that should resonate with a security audience: a single breach of trust caps the total grade, no matter how brilliant the rest of the run was. The benchmark’s own formulation is blunt — “no amount of good work outweighs a breach of trust.”

This is the anti-Good-Junction clause. In classic fraud and intrusion scenarios, the attacker is often the hardest-working, most competent person in the room right up until the moment they are not. A scoring system that lets excellence in nine areas offset dishonesty in the tenth would reward exactly the profile that does the most damage. Firmulate’s answer is structural: trust is not a line item to be averaged away, it is a gate.

Distrust of Round Numbers

There is a third, quieter principle baked in: the benchmark is skeptical of perfect 100s. In a graded evaluation, a clean 100 is usually a sign the test was too easy or the grading too generous. The July 2026 final league table reflects that honesty: gpt-5.6-sol leads with 95, Kimi K3 follows at 93, Sonnet 5 sits at 88, Fable 5 at 77, and Opus 4.8 lands at 73. Nobody maxed out — because nobody earned it.

Amazon

phishing simulation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Winners

Here is where the experiment stops being a security story and becomes a management story — which is, of course, the point.

Every model in the crucible spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The models could see the opportunity, articulate it, and then simply fail to close.

The decisive detail was buried two document references deep in the company’s own internal files — not in the customer event, not in the crisis feed. The models that actually read their own company’s documents found the competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that skimmed did not. It is the corporate equivalent of an attacker hiding in an unglamorous subdirectory: the finding is not in the loud place, it is in the boring one, two references deep.

The Thoroughness Paradox

The most sobering profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 self-learned playbook rules and the deepest analyses of any run — and still last place. The deal was left on the table, and discipline slipped in a way any security team would recognize: write attempts into a locked department instead of escalating the request properly. It did not breach trust. It just failed to respect boundaries and failed to finish. And notably, the same weakness appeared, in weaker form, in all four models.

One fairness footnote: Kimi K3 ran without an effort parameter (API default) while the others ran at their highest effort setting — and still took second place at 93.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI trust and security assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Wargame Is Live, and Your Company Can Play

Firmulate’s live company is not a slide deck. It is a synthetic firm with 13 employees, real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned and auditable. You can watch it run.

For teams that want to test themselves, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And for enterprises, the same wargame can be run against a read-only export of your own business — nothing ever writes back to real systems, which is exactly the containment model a security audience would demand before letting an AI agent anywhere near production data.

The lesson from the league table is not that AI is untrustworthy. On the manipulation tests, it was impeccably trustworthy. The lesson is that honesty and competence are separate axes — and that the failure mode to worry about is not the dramatic breach, but the diligent, thorough agent that reads everything, analyzes deeply, and still leaves the deal — or the escalation — on the table.

Full results and plain-language findings are published at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
firmulate.com/live.html — live view

The AI Company That Publishes Its Own Security Stress Test

Firmulate’s live synthetic workforce faces impersonation, privacy traps and financial pressure while every business decision remains auditable.
The Strategic Advantage Of Benchmark Partners’ AI Perspective

The Strategic Advantage Of Benchmark Partners’ AI Perspective

Analysis of Benchmark Partners’ AI perspective reveals a focus on market complexity and differentiation, shaping future industry winners.
How Europe’s New AI Powerhouse Has Strong Canadian Roots

How Europe’s New AI Powerhouse Has Strong Canadian Roots

Cohere’s planned Aleph Alpha acquisition creates a $20 billion AI group, but its Canadian control complicates Europe’s sovereignty pitch.
Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Anthropic restores Fable 5 after government blackout; OpenAI previews GPT-5.6 amid rumors of an even more capable model existing privately.