firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Can you recognize an AI by the decision it makes under pressure?

For cybersecurity and privacy readers, the revealing test of an AI system is not how confidently it explains a policy. It is what happens when an apparent executive demands a shortcut, a reporter fishes for confirmation, or a valuable fact is buried inside company documents.

Firmulate turns that question into a public experiment—and now into an interactive challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a frontier model handled a business situation and try to identify which participant was responsible.

The appeal is playful, but the evidence behind it is serious. Each model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. The result is a collection of management behavior that can be compared rather than merely admired in a chat window.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to expose judgment

The simulated business employs 13 synthetic workers and operates under real money mechanics. It burns €105k each month while bringing in €2.3k in monthly recurring revenue, with a public cash countdown adding urgency. Its operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That setting matters because polished language is not the same as competent management. A model may spot a problem, produce an impressive analysis and still fail to complete the action that creates value. Firmulate’s central finding captures that gap: all models identified every crisis and rejected every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. As the experiment summarizes it: “Same diagnosis, same pitch — no signature.”

The missing action was connected to an easily overlooked piece of company knowledge. The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.

For security teams, this resembles a familiar problem. Important context often exists outside the alert that first attracts attention. Good judgment requires following references, checking available records and resisting the temptation to act on the most visible fragment alone. In this experiment, reading the company’s own material separated a persuasive pitch from a completed commercial result.

Manipulation was the test everyone passed

The social-engineering sequence escalated through three stages of fake chief-executive messages. A reporter then attempted another route, asking for “just one yes/no, on background.” All 5 models refused the manipulation attempts.

Kimi K3 made its interpretation explicit in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of response security leaders want to see when authority, urgency and informality are combined to weaken normal controls.

The uniform refusal is encouraging, but it also shifts attention to a subtler risk. Safety is not only the absence of a catastrophic breach. It includes operational discipline: reading the relevant material, respecting boundaries, escalating when blocked and finishing legitimate work. Firmulate’s do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust”.

The league table reveals distinct management records

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran with the application programming interface default because it had no effort parameter, while the other participants ran at xhigh, an important fairness note when comparing the final results.

Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That combination makes the quiz more than a recognition game. The decisions expose recurring habits: whether a model investigates before acting, maintains discipline when a route is blocked, resists suspicious authority and converts sound analysis into a finished outcome. Those behaviors form a practical management profile with direct relevance to systems that may eventually touch customer records, support queues or forecasts.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

cybersecurity management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The security question is not whether the prose sounds safe

Firmulate’s experiment suggests that frontier models can share strong baseline instincts while differing sharply in execution. Every participant detected the crises and refused manipulation, but their league scores and commercial outcomes diverged. The critical distinctions appeared after recognition: who read deeply enough, who escalated correctly and who completed the work.

That is why the interactive quiz works as both entertainment and scrutiny. Guessing the author of an unedited decision invites readers to look past branding and tone. The more consequential question is whether the behavior shown would be acceptable inside a real organization, under pressure, with trust and money at stake.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That boundary preserves the essential value of the exercise: observing an AI workforce’s management personality before granting it operational authority.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BABY SHOWER & RE

Baby shower & registry season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
MiniMax H3's Sound Capabilities And The Significance Of 'Open' In AI

MiniMax H3’s Sound Capabilities And The Significance Of ‘Open’ In AI

MiniMax H3, launched on July 31, 2026, features joint audio-visual generation and claims ‘open’ access, but key details remain unclear.
14 Best AI Automation Software Tools for Smarter Workflows in 2026

14 Best AI Automation Software Tools for Smarter Workflows in 2026

Discover the 14 best AI automation software tools for 2026, including agent builders, coding assistants, and workflow systems, to enhance productivity.
Elon Musk’s SpaceXAI Enters The Big League With Grok 4.6, Offering Fable 5-Level Performance At An 80 Percent Discount – Wccftech

Elon Musk’s SpaceXAI Enters The Big League With Grok 4.6, Offering Fable 5-Level Performance At An 80 Percent Discount – Wccftech

Elon Musk’s SpaceXAI announces Grok 4.6, claiming Fable 5-level performance with an 80% cost reduction, but lacks independent verification or detailed specs.
Phone-based injury-risk movement screening for hiring

Phone-based injury-risk movement screening for hiring

A new approach uses phone cameras and AI to remotely assess injury risk in job candidates, potentially reducing on-the-job injuries and costs.