firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A live test of whether AI can be trusted under pressure

For cybersecurity and privacy readers, the most revealing part of Firmulate is not that its workforce is synthetic. It is that the company exposes those workers to impersonation, approval-bypass attempts, confidential-data traps and commercial pressure—and makes their decisions auditable.

Firmulate operates a small software company with 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, publishes a cash countdown and versions every workday. Its employees have accumulated more than 680 self-learned playbook rules. The result is an unusually public experiment in whether AI can run a business without surrendering discipline when urgency, authority and money collide.

This is not a staged corporate demo. The company runs every business day, and the struggle can be followed on the Firmulate live page.

Amazon

AI security stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week was designed to test trust

In the Crucible League, each frontier model received the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The final July 2026 table placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counted.

But performance was constrained by a harder principle: a single breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.” That makes the exercise particularly relevant to organizations considering AI access to customer records, support conversations, forecasts or internal documents. Competence mattered, but safe conduct was not treated as a bonus feature.

The models recognized the attacks

All five models identified every crisis and rejected every manipulation attempt. The social-engineering sequence included fake CEO messages escalating through three stages and a reporter seeking “just one yes/no, on background.” Every participant refused.

Kimi K3’s recorded reasoning captured the correct security posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because the attack relied on familiar pressures rather than an exotic technical exploit: apparent executive authority, urgency and a request framed as small enough to seem harmless.

The refusals show that frontier models can recognize common manipulation patterns in a realistic business context. Yet the wider experiment also demonstrates why refusing malicious instructions is only part of the job. A safe AI worker must protect trust while still completing legitimate work.

Security discipline did not guarantee commercial execution

Every model spotted the crises, and every model refused the manipulations. Even so, only two signed the €55,000 deal their own work had justified. Firmulate summarizes the gap as: “Same diagnosis, same pitch — no signature.”

The decisive information was not presented directly in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the references found it and won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding connects security discipline with a less discussed operational risk. An AI system may avoid an obvious breach yet still fail because it does not inspect the authorized information already available to it, or because it stops before completing the final action. In a security operations setting, the analogous failure would be recognizing an incident, drafting the right response and then neglecting the permitted step that contains the damage.

The most thorough participant still finished last

Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

This is a useful warning against treating verbosity, policy accumulation or analytical depth as proof of dependable execution. A model can understand a problem extensively and still mishandle boundaries or fail to finish. Readers can inspect the company’s public dialogue on the Firmulate quotes page, rather than relying only on polished summaries.

There is also an important fairness note: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its result should be read with that difference in mind.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Build in public becomes continuous accountability

Firmulate turns an AI-company experiment into a running security and management story. The public cash countdown supplies real pressure. Versioned workdays preserve a record. The learned rules show how behavior changes, while the league results expose the difference between noticing, reasoning and completing.

The strongest finding is not that the synthetic employees resisted fake executives and a reporter’s trick, though they did. It is that trustworthy operation requires several qualities at once: refusing manipulation, respecting access boundaries, reading authorized evidence deeply enough and carrying legitimate work through to completion.

For businesses evaluating AI workers, a polished conversation is weak evidence. Firmulate’s live company offers something harder to dismiss: visible decisions made under financial pressure, with both restraint and follow-through exposed to public scrutiny.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

cybersecurity AI model testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

Anthropic reveals that effective AI Skills are structured as folders containing instructions, scripts, and assets, transforming organizational workflows.
Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Anthropic restores Fable 5 after government blackout; OpenAI previews GPT-5.6 amid rumors of an even more capable model existing privately.
AI output review queue for customer support macros

AI output review queue for customer support macros

Support teams are testing a new AI macro review queue to ensure policy, tone, and accuracy before publishing support responses.
$965B and Climbing: Anthropic’s Series H Is Really a Compute Bet

$965B and Climbing: Anthropic’s Series H Is Really a Compute Bet

Anthropic closed a $65 billion Series H funding round at a $965 billion valuation, emphasizing compute capacity over valuation growth, signaling a major industry shift.