firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The quiet risk behind capable AI agents

Cybersecurity teams are trained to notice dramatic failures: leaked records, compromised credentials and fraudulent instructions. Firmulate’s latest experiment highlights a quieter business risk. An AI agent can identify a crisis, resist manipulation and produce persuasive work—yet still fail because it did not follow the evidence through the company’s own documents.

The decisive information in this case was a competitor weakness buried two document references deep. It was absent from the customer event that triggered the work. Models that found the file signed a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost the opportunity automatically.

That makes “reads your files before answering” more than a product promise. In a controlled business exercise, it became a measurable property with a direct commercial consequence.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model saw the crisis, but seeing was not enough

Firmulate runs frontier models as the management team of the same small software company. Each participant faced the same customers, crises and temptations during the company’s worst week. Every decision was versioned and auditable.

The company itself is deliberately unforgiving: 13 synthetic employees, burn of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 playbook rules learned through experience. The live experiment is real and watchable, allowing observers to follow how models behave when a polished answer is not the same thing as completed work.

The striking result was not a difference in basic comprehension. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. Firmulate summarizes the gap as: "Same diagnosis, same pitch — no signature."

The models that succeeded had followed the documentary trail far enough to uncover the buried competitor weakness. That information supported the full-price close. The others could understand the customer situation and construct the pitch, but their work stopped before the transaction was completed.

File-reading is becoming a security and performance question

For readers concerned with cybersecurity, privacy and corporate espionage, the result cuts in two directions. An agent must resist unauthorized requests, but it must also use legitimate internal information carefully and thoroughly. Refusing suspicious instructions is essential; failing to consult relevant authorized material can still produce an expensive operational failure.

Firmulate tested the defensive side with fake CEO messages that escalated over three stages, followed by a reporter seeking "just one yes/no, on background." All 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

This matters because social engineering often depends on urgency, authority and requests that appear harmless in isolation. In Firmulate’s exercise, the models did not take the bait. The more revealing separation came afterward: whether they could navigate trusted company knowledge and carry an approved business process to its conclusion.

The league rewards completion and trust

In the final July 2026 Crucible League, gpt-5.6-sol ranked first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts. A single breach of trust caps the total under the principle that "no amount of good work outweighs a breach of trust." The complete standings and findings are available on Firmulate’s public benchmark page.

K3’s result also carries an important fairness note: it ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it is relevant context when comparing placements.

Opus 4.8 illustrates why thoroughness alone is an unreliable proxy for effectiveness. It produced the deepest analyses and learned 80 additional rules, yet finished last. The close was left on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. That record invites readers to judge behavior rather than reputation: which model sounds decisive, which verifies its evidence, and which actually finishes the work?

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Enterprise MCP Security: Securing AI Agents, Tools & LLM Operations

Enterprise MCP Security: Securing AI Agents, Tools & LLM Operations

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should measure before granting access

The experiment suggests that evaluating an AI workforce requires more than testing whether a model writes clearly or recognizes an obvious attack. Buyers need evidence that an agent follows references, locates relevant authorized information, respects access boundaries, escalates when blocked and completes the business action its reasoning supports.

Firmulate offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That separation is especially relevant for organizations balancing useful access against privacy, security and control.

The buried fact is the larger lesson. Important corporate knowledge rarely arrives neatly inside the prompt. It lives in files, linked references and institutional rules. An agent that does not inspect that context may still look intelligent. In this experiment, however, the difference between looking capable and delivering value was a signed €55,000 deal at full price.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and privacy solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
When AI Builds Itself: Inside Anthropic’s Evidence on Recursive Self-Improvement

When AI Builds Itself: Inside Anthropic’s Evidence on Recursive Self-Improvement

Anthropic presents data indicating AI systems are increasingly capable of automating AI research tasks, raising the possibility of recursive self-improvement.
Capital: The Lever Beneath the Levers

Capital: The Lever Beneath the Levers

Exploring how capital funding shapes AI development, with recent IPOs and circular investments revealing vulnerabilities in the industry’s financial structure.
Should You Use Mistral Forge? A Buyer’s Decision Guide

Should You Use Mistral Forge? A Buyer’s Decision Guide

A detailed analysis of Mistral Forge’s suitability for enterprise AI, covering who it fits, alternatives, and red flags to watch for.
How Europe’s New AI Powerhouse Has Strong Canadian Roots

How Europe’s New AI Powerhouse Has Strong Canadian Roots

Cohere’s planned Aleph Alpha acquisition creates a $20 billion AI group, but its Canadian control complicates Europe’s sovereignty pitch.