Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows the lowest-performing AI scores 26 points out of a possible high score, emphasizing the value of partial work and trust. The test reveals how AI models handle crises, documentation, and trust breaches.

A recent AI management benchmark conducted by Firmulate reveals that even the poorest AI managers score 26 points out of a possible 100, not zero, in a test designed to simulate a company’s worst week. This benchmark is discussed in the original analysis. This scoring system emphasizes partial progress and trust, with the highest scorer achieving 95 points. The results challenge assumptions about AI performance metrics and highlight the importance of integrity in automated management systems.

The benchmark involved four frontier AI models managing a small business through seven days of crises, including customer issues and social engineering attacks. Each model was evaluated on its ability to handle real-world management tasks, with scores reflecting both effectiveness and trustworthiness. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort—doing almost nothing—earned 26 points, underscoring that partial work is valued in this system.

The scoring system is designed to discourage dishonest or overly perfect performance. A perfect score of 100 is intentionally avoided, as it would suggest unmeasured or suspiciously flawless management. The benchmark’s principle is clear: “no amount of good work outweighs a breach of trust,” highlighting the importance of integrity in AI management systems. A single breach, such as failing to escalate or read documentation, caps the score at that level, regardless of other successes. This approach aims to mirror real-world expectations where integrity is paramount.

Interestingly, models that read their own documentation and verified critical information secured full deals, earning higher points. For example, two models identified key references deep in the company’s files and closed a €55,000 deal, demonstrating the importance of thoroughness and accurate information retrieval. Conversely, models that failed to do so missed out on significant revenue, showing that detailed reading and follow-through are crucial skills for AI managers in high-stakes environments.

At a glance
reportWhen: published July 2026
The developmentA new benchmark testing AI management under stress shows the lowest score is 26 points, not zero, raising questions about how AI effectiveness is measured.
Why The Worst AI Manager Still Gets 26 Points
AI Management Benchmark · Firmulate · July 2026

Why the Worst AI Manager Still Gets 26 Points

Inside a benchmark that refuses to hand out zeros: four frontier AI models managed a small business through seven days of crises — customer meltdowns, social engineering attacks, and a €55,000 sales negotiation. The lowest possible score isn’t zero, and the highest isn’t 100. Here’s why.

“No amount of good work outweighs a breach of trust.”

Core scoring principle of the benchmark
26
Baseline score for a do-nothing manager
95
Top score — gpt-5.6-sol
4
Frontier AI models tested
7
Days of simulated company crises
01

The Scoreboard: No Zeros, No Hundreds

ModelScore / 100Read Docs & VerifiedClosed €55K DealTrust Breach
gpt-5.6-sol95None
Opus 4.873None
Do-nothing baseline26None
Perfect score100✗ Explicitly disallowed — flagged as unmeasured or suspiciously flawless
Scoring Rule · Floor

Why 26, Not Zero

Even a manager who does almost nothing earns 26 points for being present in the seat. The benchmark values minimal effort and recognizes the simplest management actions — partial work counts.

Scoring Rule · Ceiling

Why Never 100

A perfect score is intentionally withheld. Flawless management is treated as either unmeasurable or a sign of gaming — the system discourages suspiciously perfect performance.

Scoring Rule · Cap

The Breach Cap

One trust breach — failing to escalate, skipping documentation, not verifying critical information — caps the score at that level, no matter how brilliant every other decision was.

02

Score Spectrum: From Baseline to Near-Perfect

gpt-5.6-sol — top performer95
Opus 4.8 — lowest model73
Do-nothing baseline26

The €55,000 lesson: two models dug deep into the company’s files, identified key references buried in documentation, verified critical information, and closed a €55,000 deal in full. Models that skipped the reading missed the revenue entirely — thoroughness pays, literally.

03

Anatomy of the Worst Week

1

📄 Read the Docs

Models dig through company files to find buried references — the difference between winning and losing the deal.

2

🚨 Triage Crises

Seven days of customer meltdowns and operational fires test prioritization under sustained pressure.

3

🛡️ Defend Trust

Social engineering attacks probe whether models escalate suspicious requests or quietly comply.

4

🤝 Close the Deal

Accurate information retrieval and follow-through determine whether the €55,000 deal closes in full.

04

Key Questions, Answered

Q1Why doesn’t the lowest score drop to zero?

The 26-point baseline reflects minimal effort: the scoring system recognizes that partial work has value, crediting even the simplest management actions.

Q2What would a perfect score of 100 mean?

Nothing — it’s deliberately impossible. Perfect management is treated as unmeasured or suspiciously flawless, so the benchmark never awards it.

Q3How do trust breaches affect scores?

A single breach — failing to escalate or verify information — caps the score at that level regardless of every other success. Integrity outweighs competence.

Q4Can this predict real-world enterprise performance?

Not yet. Live validation is pending via ongoing simulations at firmulate.com/live, which model real business mechanics with synthetic employees.

Implications of Partial Progress and Trust in AI Management

This benchmark underscores that in AI management, partial progress—such as triaging crises, reading documentation, and maintaining trust—is highly valued. It shifts focus from perfect performance to consistent, trustworthy effort, which is vital for deploying AI in real business contexts. The emphasis on trust breaches as a scoring cap highlights the importance of integrity over mere competence, especially when AI controls sensitive operations like customer support or sales.

For organizations considering AI management tools, the results suggest that effectiveness depends not only on how well models communicate but also on their ability to follow through, verify information, and act ethically under pressure. The benchmark’s transparent, auditable scoring system offers a method for evaluating AI suitability for critical tasks, potentially guiding enterprise adoption and risk management strategies.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Firmulate Benchmark and Its Design Principles

Developed by Firmulate, the benchmark aims to evaluate AI models in realistic management scenarios, focusing on their ability to handle crises, verify information, and maintain trust. Unlike traditional benchmarks that measure language proficiency or narrow tasks, this league simulates a company’s worst week, with models managing customer issues, social engineering threats, and sales negotiations. The scoring system intentionally avoids zeros and perfect scores, instead rewarding partial work and penalizing breaches of trust.

The benchmark’s design reflects a core principle: in real business, partial progress and integrity are more valuable than flawless but untrustworthy performance. The scoring system caps the total score at 26 for minimal effort, with the highest scores approaching 95, and explicitly disallows perfect 100s to prevent gaming or unmeasured performance. This approach aims to foster more realistic, trustworthy AI management solutions.

“A manager who does nothing still earns 26 points, emphasizing that partial work and trustworthiness are core to this benchmark.”

— an anonymous researcher

Amazon

AI documentation reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Benchmark Scope and Real-World Applicability

It is not yet clear how these benchmark results will translate to real-world enterprise environments beyond simulated crises. The models’ performance in managing actual business operations, handling complex negotiations, or maintaining long-term trust remains to be tested in live settings. Additionally, the impact of different business sizes, industries, or management styles on AI effectiveness is still unknown.

Furthermore, the scoring system’s emphasis on trust breaches may not capture all dimensions of AI management quality, such as innovation or strategic thinking. The long-term implications of prioritizing integrity over performance are still under discussion among AI developers and business leaders.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Evaluating AI Management in Business Settings

Organizations interested in deploying AI management tools can engage with the benchmark through live experiments, such as the ongoing simulation at firmulate.com/live, which models real business mechanics with synthetic employees. Future developments may include expanding the benchmark to different industries, testing more models, and refining scoring criteria to better reflect real-world complexities.

Researchers and developers are likely to analyze the detailed decision logs to improve AI models’ ability to verify information and maintain trust. The results may also influence industry standards for AI transparency, accountability, and trustworthiness, shaping how AI is integrated into enterprise operations over the coming months.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the lowest score in the benchmark not drop to zero?

The benchmark assigns 26 points to the do-nothing baseline to reflect minimal effort, emphasizing that partial work has value and that the scoring system recognizes even the simplest management efforts.

What does a perfect score of 100 indicate in this benchmark?

The benchmark explicitly avoids assigning a 100 to prevent unmeasured or suspiciously flawless performance, indicating that perfect management is either unmeasurable or undesirable in this context.

How do trust breaches affect AI scores in the benchmark?

A single breach of trust, such as failing to escalate or verify information, caps the score at that level regardless of other successes, underscoring the importance of integrity in AI management.

Can this benchmark predict real-world AI performance in enterprises?

While it provides valuable insights into AI handling crises and trust, its direct applicability to real-world operations remains to be validated through live testing and industry adoption.

What should companies consider before deploying AI managers based on these results?

They should evaluate whether the AI can verify critical information, follow through on tasks, and maintain trust under pressure, as these are key factors highlighted by the benchmark.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
AI Is the Alibi. The Reorg Is the Signal.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs and restructuring are linked to AI ambitions, but evidence suggests market pressures and crypto downturns are the primary causes.
The City That Watches Itself: The Living Digital Twin, And The God’s-Eye View We’re Building

The City That Watches Itself: The Living Digital Twin, And The God’s-Eye View We’re Building

Cities are developing dynamic digital twins that combine real-time data and AI to improve planning and monitoring, raising both opportunities and privacy concerns.
Outcome-First Decisions: The Friction Is The Feature

Outcome-First Decisions: The Friction Is The Feature

A new decision framework prioritizes decisive, testable choices over plans, reducing wasted effort and building decision calibration over time.
AI Scope-of-work Reviewer For Agency Selection

AI Scope-of-work Reviewer For Agency Selection

AI is being tested as a scope-of-work reviewer to assist SMBs and mid-market companies in evaluating marketing agency proposals more effectively.