
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Can you recognize an AI by the decision it makes under pressure?
For cybersecurity and privacy readers, the revealing test of an AI system is not how confidently it explains a policy. It is what happens when an apparent executive demands a shortcut, a reporter fishes for confirmation, or a valuable fact is buried inside company documents.
Firmulate turns that question into a public experiment—and now into an interactive challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a frontier model handled a business situation and try to identify which participant was responsible.
The appeal is playful, but the evidence behind it is serious. Each model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. The result is a collection of management behavior that can be compared rather than merely admired in a chat window.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company designed to expose judgment
The simulated business employs 13 synthetic workers and operates under real money mechanics. It burns €105k each month while bringing in €2.3k in monthly recurring revenue, with a public cash countdown adding urgency. Its operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
That setting matters because polished language is not the same as competent management. A model may spot a problem, produce an impressive analysis and still fail to complete the action that creates value. Firmulate’s central finding captures that gap: all models identified every crisis and rejected every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. As the experiment summarizes it: “Same diagnosis, same pitch — no signature.”
The missing action was connected to an easily overlooked piece of company knowledge. The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
For security teams, this resembles a familiar problem. Important context often exists outside the alert that first attracts attention. Good judgment requires following references, checking available records and resisting the temptation to act on the most visible fragment alone. In this experiment, reading the company’s own material separated a persuasive pitch from a completed commercial result.
Manipulation was the test everyone passed
The social-engineering sequence escalated through three stages of fake chief-executive messages. A reporter then attempted another route, asking for “just one yes/no, on background.” All 5 models refused the manipulation attempts.
Kimi K3 made its interpretation explicit in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of response security leaders want to see when authority, urgency and informality are combined to weaken normal controls.
The uniform refusal is encouraging, but it also shifts attention to a subtler risk. Safety is not only the absence of a catastrophic breach. It includes operational discipline: reading the relevant material, respecting boundaries, escalating when blocked and finishing legitimate work. Firmulate’s do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust”.
The league table reveals distinct management records
The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran with the application programming interface default because it had no effort parameter, while the other participants ran at xhigh, an important fairness note when comparing the final results.
Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
That combination makes the quiz more than a recognition game. The decisions expose recurring habits: whether a model investigates before acting, maintains discipline when a route is blocked, resists suspicious authority and converts sound analysis into a finished outcome. Those behaviors form a practical management profile with direct relevance to systems that may eventually touch customer records, support queues or forecasts.

cybersecurity management decision tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The security question is not whether the prose sounds safe
Firmulate’s experiment suggests that frontier models can share strong baseline instincts while differing sharply in execution. Every participant detected the crises and refused manipulation, but their league scores and commercial outcomes diverged. The critical distinctions appeared after recognition: who read deeply enough, who escalated correctly and who completed the work.
That is why the interactive quiz works as both entertainment and scrutiny. Guessing the author of an unedited decision invites readers to look past branding and tone. The more consequential question is whether the behavior shown would be acceptable inside a real organization, under pressure, with trust and money at stake.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That boundary preserves the essential value of the exercise: observing an AI workforce’s management personality before granting it operational authority.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.