📊 Full opportunity report: The Management Test That Offers Insights Into AI’s Genuine Work Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
Create a free accountAs an affiliate, we earn on qualifying purchases.
TL;DR
Firmulate has launched a live management test using real management decisions to evaluate AI models’ ability to handle business crises. The experiment reveals significant differences in AI performance, especially in action and trustworthiness, not just analysis. This could influence how enterprises evaluate AI automation for management tasks.
Firmulate has launched a live management test that evaluates AI models on their ability to handle a simulated company’s worst week, providing insights into their decision-making, trustworthiness, and execution. This experiment aims to assess model performance under realistic conditions, moving beyond theoretical analysis or polished responses. You can learn more about this approach in the detailed report.
The experiment involves five AI models managing a small software company facing crises, with decisions based on 242 real, unedited management actions. For more on evaluating AI decision-making, see the original analysis. The models are scored on their ability to diagnose problems, signal trustworthiness, follow through, and execute critical tasks such as closing deals. The results, published in July 2026, rank GPT-5.6-SOL first with 95 points, followed by Kimi K3, Sonnet 5, Fable 5, and Opus 4.8, with the baseline scoring only 26 points. This kind of evaluation is similar to the management test that exposes an AI’s real working style.
All models identified crises and refused manipulative attempts, such as fake CEO messages. However, only two models successfully completed a key €55,000 deal, despite all recognizing the opportunity. The experiment highlights that effective analysis alone does not guarantee successful management—execution and trust are also essential. For instance, Opus 4.8, despite thorough analysis, failed to close a deal due to operational issues.
The test that exposes AI’s genuine work style
Firmulate placed five AI models inside a simulated software company’s worst week. The result: recognizing a crisis was common. Acting decisively, earning trust and finishing critical work were not.
Real, unedited management actions shaped the simulation.
GPT-5.6-SOL ranked first against a 26-point baseline.
Only two models completed the crucial €55,000 deal.
Management is more than a correct diagnosis
Traditional benchmarks reward polished answers. This live test follows what happens after the answer—when an agent must navigate pressure, people, uncertainty and operational friction.
Diagnose
Identify emerging crises, distinguish symptoms from root causes and prioritize the real threat.
Build trust
Resist manipulation, communicate honestly and avoid actions that undermine organizational confidence.
Follow through
Preserve context across steps, maintain commitments and keep important work moving under pressure.
Execute
Translate insight into completed tasks, including the decisive act of closing a commercial deal.
Performance spread reveals distinct work styles
The top score reached 95 points, far beyond the 26-point baseline. The ordering reflects combined management performance—not analytical eloquence alone.
Exact point totals were supplied for the winner and baseline. Intermediate bars visualize published rank order, not undisclosed scores.
Knowing what to do did not mean getting it done
Every model recognized danger and rejected fake CEO manipulation. The decisive separation appeared in operational completion.
| Observed capability | All models | Top performers | Management signal |
|---|---|---|---|
| Crisis identification | ✓ Successful | ✓ Successful | Analysis was broadly strong |
| Fake CEO resistance | ✓ Refused | ✓ Refused | Basic trust safeguards held |
| €55,000 opportunity recognized | ✓ Recognized | ✓ Recognized | Commercial awareness was present |
| €55,000 deal completed | 2/5 Completed | ✓ Executed | Follow-through became the separator |
| Operational consistency | ~ Variable | ✓ Stronger | Reliable action mattered more than verbosity |
Case in point: Opus 4.8 produced thorough analysis but failed to close the deal because of operational problems.
How a live test exposes real working behavior
Each stage creates evidence that an enterprise can inspect—from the first signal to the final business result.
Crisis appears
Conflicting priorities, manipulation attempts and commercial risk enter the environment.
Model decides
The AI diagnoses the situation, chooses priorities and signals how it handles trust.
Actions unfold
Plans encounter tools, dependencies, interruptions and the need for sustained follow-through.
Outcome is scored
Completed work reveals operational discipline that a polished answer cannot demonstrate.
“Operational discipline and trustworthiness are as important as analytical depth.”
Core finding from the experimentEvaluate AI in realistic, instrumented scenarios before automating consequential management work.
Measure decisions, tool use, recovery and completion—not only the quality of a model’s explanation.
Test across more industries, larger organizations and longer operational time horizons.
The unanswered management questions
The experiment is a strong operational signal, but its small-company crisis setting does not establish universal performance.
Will performance hold in larger organizations?
More stakeholders, approvals and dependencies could expose different failure modes.
Will the ranking survive other industries?
Regulated, physical and safety-critical environments may demand different capabilities.
Can operational discipline persist over time?
Long-running work may test memory, consistency, escalation and strategic coherence.
The most useful AI management test is not “Can the model explain the right move?” It is “Can the model make the move, preserve trust and deliver the outcome?”
Implications for AI in Business Management
This experiment suggests that the ability to analyze problems is not sufficient for effective management; operational execution, trust management, and follow-through are important factors. For organizations considering AI automation, testing models in realistic scenarios can provide a clearer picture of their practical management capabilities. The findings challenge the assumption that more analysis or thoroughness automatically leads to better outcomes, emphasizing the importance of operational discipline in AI performance.

128GB Flash Drive Aiibe USB Flash Drive 128 GB Thumb Drive USB 2.0 Memory Stick Zip Drive Backup Jump Drive Single 128GB 128G USB Drive for PC Laptop
- Large Storage Capacity: 128GB for files, photos, videos, music
- Plug and Play: No software needed, easy to use
- Wide Device Compatibility: Works with PC, Mac, TV, car, and more
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing
Traditional AI benchmarks tend to focus on analysis, language understanding, or problem-solving in controlled settings. However, real-world management involves decision-making under pressure, establishing trust, and executing plans, which are less frequently tested. Firmulate’s experiment builds on recent efforts to evaluate AI in operational contexts, using live decision-making scenarios modeled after actual business crises. The results published in July 2026 represent a step toward understanding AI’s potential in practical management roles.
“Testing AI models against real management decisions provides insights into their practical capabilities and limitations beyond theoretical analysis.”
— Firmulate spokesperson
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About AI Management Performance
It remains to be seen how these results will translate to different industries or larger organizations. The experiment focused on a small software company under specific crisis conditions, and performance may vary in other scenarios or more complex environments. Additionally, the long-term implications of deploying AI models with these capabilities in live business settings are still under investigation, including potential risks related to trust, operational errors, and strategic decision-making.

CloudValley Laptop Camera Cover Slide, Metal 0.023 Inch Ultra-Thin, 2 Packs
- Privacy Protection: Ensures privacy on laptops and tablets
- Fashionable Design: Elegant space aluminum alloy finish
- Ultra-Thin Profile: Only 0.023 inch thick for seamless use
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation
Firmulate plans to expand its testing framework to include more diverse scenarios and larger organizations. Further research will explore how AI models develop operational discipline over time and how they can be integrated into ongoing management processes. Organizations are encouraged to conduct internal live tests to assess AI readiness before full deployment, using insights gained from this experiment to improve evaluation methods.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is this management test significant for AI development?
This test offers insights into how AI models perform in real management scenarios, focusing on decision-making, trust, and execution—areas that are critical for operational success and are often underrepresented in traditional benchmarks.
What does the experiment reveal about AI’s ability to handle crises?
The experiment indicates that while AI models can identify crises and resist manipulative tactics, their ability to follow through and successfully close deals varies, highlighting areas for potential improvement in operational consistency.
Could this testing approach be used by businesses to evaluate their own AI tools?
Yes, organizations can implement similar live scenarios with their data to assess the practical management capabilities of their AI models prior to broader deployment.
What are the limitations of this experiment?
The experiment is limited to a specific industry and crisis type; results may differ in other contexts. Further research is needed to understand broader applicability and long-term effects.
What is the main takeaway for AI developers and users?
Effective management with AI involves more than analysis; operational discipline, trustworthiness, and follow-through are key factors for success in real-world applications.
Source: ThorstenMeyerAI.com
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.