The Management Test That Offers Insights Into AI’s Genuine Work Style
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Offers Insights Into AI’s Genuine Work Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate has launched a live management test using real management decisions to evaluate AI models’ ability to handle business crises. The experiment reveals significant differences in AI performance, especially in action and trustworthiness, not just analysis. This could influence how enterprises evaluate AI automation for management tasks.

Firmulate has launched a live management test that evaluates AI models on their ability to handle a simulated company’s worst week, providing insights into their decision-making, trustworthiness, and execution. This experiment aims to assess model performance under realistic conditions, moving beyond theoretical analysis or polished responses. You can learn more about this approach in the detailed report.

The experiment involves five AI models managing a small software company facing crises, with decisions based on 242 real, unedited management actions. For more on evaluating AI decision-making, see the original analysis. The models are scored on their ability to diagnose problems, signal trustworthiness, follow through, and execute critical tasks such as closing deals. The results, published in July 2026, rank GPT-5.6-SOL first with 95 points, followed by Kimi K3, Sonnet 5, Fable 5, and Opus 4.8, with the baseline scoring only 26 points. This kind of evaluation is similar to the management test that exposes an AI’s real working style.

All models identified crises and refused manipulative attempts, such as fake CEO messages. However, only two models successfully completed a key €55,000 deal, despite all recognizing the opportunity. The experiment highlights that effective analysis alone does not guarantee successful management—execution and trust are also essential. For instance, Opus 4.8, despite thorough analysis, failed to close a deal due to operational issues.

At a glance
reportWhen: ongoing, with results published in July…
The developmentFirmulate’s live experiment tests AI models on managing a simulated company crisis, revealing their decision-making and operational capabilities.
The Management Test That Offers Insights Into AI’s Genuine Work Style
Live AI management benchmark · July 2026

The test that exposes AI’s genuine work style

Firmulate placed five AI models inside a simulated software company’s worst week. The result: recognizing a crisis was common. Acting decisively, earning trust and finishing critical work were not.

Decisions in the test 242

Real, unedited management actions shaped the simulation.

Winning score 95 pts

GPT-5.6-SOL ranked first against a 26-point baseline.

Deal execution 2 of 5

Only two models completed the crucial €55,000 deal.

5 AI models tested
€55K Critical deal value
100% Detected the crises
40% Closed the deal
What the benchmark measures

Management is more than a correct diagnosis

Traditional benchmarks reward polished answers. This live test follows what happens after the answer—when an agent must navigate pressure, people, uncertainty and operational friction.

01 Sensemaking

Diagnose

Identify emerging crises, distinguish symptoms from root causes and prioritize the real threat.

02 Integrity

Build trust

Resist manipulation, communicate honestly and avoid actions that undermine organizational confidence.

03 Continuity

Follow through

Preserve context across steps, maintain commitments and keep important work moving under pressure.

04 Operations

Execute

Translate insight into completed tasks, including the decisive act of closing a commercial deal.

Published ranking

Performance spread reveals distinct work styles

The top score reached 95 points, far beyond the 26-point baseline. The ordering reflects combined management performance—not analytical eloquence alone.

GPT-5.6-SOL
95
Kimi K3
#2
Sonnet 5
#3
Fable 5
#4
Opus 4.8
#5
Baseline
26
Score scale
0 50 100
Pts

Exact point totals were supplied for the winner and baseline. Intermediate bars visualize published rank order, not undisclosed scores.

Analysis versus action

Knowing what to do did not mean getting it done

Every model recognized danger and rejected fake CEO manipulation. The decisive separation appeared in operational completion.

Observed capability All models Top performers Management signal
Crisis identification Successful Successful Analysis was broadly strong
Fake CEO resistance Refused Refused Basic trust safeguards held
€55,000 opportunity recognized Recognized Recognized Commercial awareness was present
€55,000 deal completed 2/5 Completed Executed Follow-through became the separator
Operational consistency ~ Variable Stronger Reliable action mattered more than verbosity

Case in point: Opus 4.8 produced thorough analysis but failed to close the deal because of operational problems.

Traceability chain

How a live test exposes real working behavior

Each stage creates evidence that an enterprise can inspect—from the first signal to the final business result.

1 Pressure

Crisis appears

Conflicting priorities, manipulation attempts and commercial risk enter the environment.

2 Judgment

Model decides

The AI diagnoses the situation, chooses priorities and signals how it handles trust.

3 Behavior

Actions unfold

Plans encounter tools, dependencies, interruptions and the need for sustained follow-through.

4 Evidence

Outcome is scored

Completed work reveals operational discipline that a polished answer cannot demonstrate.

“Operational discipline and trustworthiness are as important as analytical depth.”

Core finding from the experiment
Enterprise implication

Evaluate AI in realistic, instrumented scenarios before automating consequential management work.

Benchmark implication

Measure decisions, tool use, recovery and completion—not only the quality of a model’s explanation.

Research implication

Test across more industries, larger organizations and longer operational time horizons.

What comes next

The unanswered management questions

The experiment is a strong operational signal, but its small-company crisis setting does not establish universal performance.

Scale

Will performance hold in larger organizations?

More stakeholders, approvals and dependencies could expose different failure modes.

Transferability

Will the ranking survive other industries?

Regulated, physical and safety-critical environments may demand different capabilities.

Duration

Can operational discipline persist over time?

Long-running work may test memory, consistency, escalation and strategic coherence.

Main takeaway

The most useful AI management test is not “Can the model explain the right move?” It is “Can the model make the move, preserve trust and deliver the outcome?”

Implications for AI in Business Management

This experiment suggests that the ability to analyze problems is not sufficient for effective management; operational execution, trust management, and follow-through are important factors. For organizations considering AI automation, testing models in realistic scenarios can provide a clearer picture of their practical management capabilities. The findings challenge the assumption that more analysis or thoroughness automatically leads to better outcomes, emphasizing the importance of operational discipline in AI performance.

128GB Flash Drive Aiibe USB Flash Drive 128 GB Thumb Drive USB 2.0 Memory Stick Zip Drive Backup Jump Drive Single 128GB 128G USB Drive for PC Laptop

128GB Flash Drive Aiibe USB Flash Drive 128 GB Thumb Drive USB 2.0 Memory Stick Zip Drive Backup Jump Drive Single 128GB 128G USB Drive for PC Laptop

  • Large Storage Capacity: 128GB for files, photos, videos, music
  • Plug and Play: No software needed, easy to use
  • Wide Device Compatibility: Works with PC, Mac, TV, car, and more

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing

Traditional AI benchmarks tend to focus on analysis, language understanding, or problem-solving in controlled settings. However, real-world management involves decision-making under pressure, establishing trust, and executing plans, which are less frequently tested. Firmulate’s experiment builds on recent efforts to evaluate AI in operational contexts, using live decision-making scenarios modeled after actual business crises. The results published in July 2026 represent a step toward understanding AI’s potential in practical management roles.

“Testing AI models against real management decisions provides insights into their practical capabilities and limitations beyond theoretical analysis.”

— Firmulate spokesperson

Amazon

laptop privacy screen protectors

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Management Performance

It remains to be seen how these results will translate to different industries or larger organizations. The experiment focused on a small software company under specific crisis conditions, and performance may vary in other scenarios or more complex environments. Additionally, the long-term implications of deploying AI models with these capabilities in live business settings are still under investigation, including potential risks related to trust, operational errors, and strategic decision-making.

CloudValley Laptop Camera Cover Slide, Metal 0.023 Inch Ultra-Thin, 2 Packs

CloudValley Laptop Camera Cover Slide, Metal 0.023 Inch Ultra-Thin, 2 Packs

  • Privacy Protection: Ensures privacy on laptops and tablets
  • Fashionable Design: Elegant space aluminum alloy finish
  • Ultra-Thin Profile: Only 0.023 inch thick for seamless use

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation

Firmulate plans to expand its testing framework to include more diverse scenarios and larger organizations. Further research will explore how AI models develop operational discipline over time and how they can be integrated into ongoing management processes. Organizations are encouraged to conduct internal live tests to assess AI readiness before full deployment, using insights gained from this experiment to improve evaluation methods.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is this management test significant for AI development?

This test offers insights into how AI models perform in real management scenarios, focusing on decision-making, trust, and execution—areas that are critical for operational success and are often underrepresented in traditional benchmarks.

What does the experiment reveal about AI’s ability to handle crises?

The experiment indicates that while AI models can identify crises and resist manipulative tactics, their ability to follow through and successfully close deals varies, highlighting areas for potential improvement in operational consistency.

Could this testing approach be used by businesses to evaluate their own AI tools?

Yes, organizations can implement similar live scenarios with their data to assess the practical management capabilities of their AI models prior to broader deployment.

What are the limitations of this experiment?

The experiment is limited to a specific industry and crisis type; results may differ in other contexts. Further research is needed to understand broader applicability and long-term effects.

What is the main takeaway for AI developers and users?

Effective management with AI involves more than analysis; operational discipline, trustworthiness, and follow-through are key factors for success in real-world applications.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
Founders Fund’s outlier bet on humanely killed fish

Founders Fund’s outlier bet on humanely killed fish

Founders Fund has backed Shinkei’s innovative approach to fish harvesting, focusing on humane killing and supply chain re-shoring, marking a new investment trend.
Can The AI4S Seed Scientist Program Turn STEM Brain Drain Around?

Can The AI4S Seed Scientist Program Turn STEM Brain Drain Around?

ByteDance launches a six-month pilot to recruit 100 scientists for AI-driven research, seeking to address scientific talent loss and boost AI4S efforts.
Austria Lobbies EU to Host Anthropic After US Access Curbs

Austria Lobbies EU to Host Anthropic After US Access Curbs

Austria is lobbying the EU to host Anthropic to counter US efforts restricting foreign access to advanced AI models.
AI’s Management Gap Appears After The Right Answer

AI’s Management Gap Appears After The Right Answer

A recent experiment reveals AI models can diagnose and respond accurately but struggle to complete trustworthy, actionable work under real-world pressures.