firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A convincing message from the boss can be a security test. So can a reporter’s promise that a question is “on background.” In Firmulate’s company simulation, five frontier models refused those approaches. The harder test came afterward: would they find a crucial fact buried in company files, then follow through on what they had learned?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same crises, different outcomes

Firmulate ran each model through the same worst week at a small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. All five models spotted every crisis and refused every manipulation attempt, including fake CEO messages escalated over three stages and the reporter’s “just one yes/no, on background” request.

Kimi K3’s on-record reasoning on the reporter approach was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear security instinct. But the results also show why resisting an obvious lure is only part of the job.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The evidence was in the files

The decisive competitor weakness was not in the customer event. It was buried two document references deep in the company’s own files. Models that read the file could win a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Yet only two models signed the deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

In the final July 2026 league table, gpt-5.6-sol placed first with 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. K3 finished ahead of three of the four Western frontier models in the field, and Firmulate describes it as the cleanest in discipline, with one deviation.

That result makes the contest harder to reduce to a single idea of intelligence. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four. Thorough analysis did not guarantee a completed decision.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A live test with real stakes inside the simulation

Firmulate presents the experiment as a watchable, ongoing company rather than a static demo. Its simulated business has 13 employees and real money mechanics: burn of €105k a month against €2.3k in monthly recurring revenue, plus a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and each workday is versioned. Readers can watch the company live or explore the benchmark and plain-language findings.

For security and privacy readers, the point is practical: an agent can refuse impersonation and still fail to inspect the evidence that matters, or fail to carry an earned conclusion through to action. Firmulate says enterprises can run the wargame against a read-only export of their own business; the pilot does not write back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before you trust

Kimi K3’s second-place finish puts it ahead of three Western frontier models in this trial, while the overall results show that catching every crisis and resisting every bait still did not ensure the job was finished. The league is open. If an AI model may touch your business, choosing one without testing it on your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI deal signing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
The Local-First Agentic Operator

The Local-First Agentic Operator

A single operator, leveraging agentic AI, now builds and manages diverse software products previously requiring entire organizations, emphasizing local-first, provider-agnostic principles.
The 12 Questions That Every AI Enthusiast Needs To Know

The 12 Questions That Every AI Enthusiast Needs To Know

A comprehensive guide to the 12 key questions about AI, explaining how it works, its limitations, and why understanding these questions matters.
Fashion Signal Monitor: Teyana Taylor Wows In Bold Burgundy Gown At The 2026 BET Awards

Fashion Signal Monitor: Teyana Taylor Wows In Bold Burgundy Gown At The 2026 BET Awards

Teyana Taylor captivates at the 2026 BET Awards wearing a bold burgundy gown, setting a new fashion signal in the industry.
The Real Cost Of A Local-Inference Rig In 2026

The Real Cost Of A Local-Inference Rig In 2026

Analyzing the true expenses of building a local AI inference setup in 2026, including hardware costs, VRAM constraints, and strategic choices for different model sizes.