The AI Company Turning Corporate Survival Into A Live Feed

📊 Full opportunity report: The AI Company Turning Corporate Survival Into A Live Feed on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Firmulate has launched a live experiment with a synthetic workforce managing a software company. The project reveals that while AI can diagnose problems, completing actions remains a critical challenge. This experiment offers new insights into AI’s practical capabilities and limitations, as detailed in the original analysis.

Firmulate has launched a live experiment where a synthetic workforce of 13 AI agents manages a software company, exposing the practical challenges of automation in real-time. This public project reveals that, despite sophisticated analysis and rules, AI systems often fail to complete critical business actions, risking company survival. The experiment’s transparency offers a rare view into AI’s operational limits and the importance of disciplined execution.

The experiment involves a synthetic team operating a company with a monthly burn rate of €105,000 against €2,300 in recurring revenue, illustrating the challenges of turning nonprofits into companies. Every workday is versioned, creating an evolving record of decisions, actions, successes, and failures. Over 680 self-learned rules have been developed, yet the project demonstrates that thorough analysis alone does not guarantee successful business outcomes. The core issue remains the AI’s ability to turn diagnosis into action, with many models recognizing problems but failing to complete critical tasks. This aligns with the insights from the original analysis.

In the latest results, only two of five AI models secured a €55,000 deal, despite all identifying the opportunity. The decisive factor was uncovering a hidden weakness in the customer’s decision process, buried in the company’s files—an insight that only some models followed through to act upon. The experiment also tested trust, with AI agents refusing fake approval requests, showing that trust is maintained through evidence retrieval and disciplined action rather than superficial compliance. The top-performing model, gpt-5.6-sol, scored 95 out of 100, while the most thorough, Opus 4.8, scored only 73, highlighting that more analysis does not necessarily lead to better results.

At a glance
breakingWhen: ongoing, with real-time updates availab…
The developmentFirmulate is running a live AI-driven company with 13 synthetic employees, exposing the gap between AI diagnosis and execution amid ongoing cash pressure.

Implications of AI’s Inability to Complete Business Actions

This experiment underscores a critical challenge for AI automation: diagnosis and recognition of problems are insufficient without reliable execution. For businesses, it highlights that AI’s value depends on its capacity to follow through on decisions, resist pressure, and respect organizational boundaries. The live, public nature of the experiment makes these limitations visible, offering a new lens for evaluating AI tools beyond their analytical capabilities. It also raises questions about the future of AI-driven management and the importance of disciplined, verifiable actions for organizational survival.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Live Automation Experiments

Firmulate’s project builds on the trend of transparency in AI development, where companies share their operational experiments publicly. Unlike typical demos or product launches, this ongoing experiment involves a synthetic workforce managing a real company’s decisions and cash flow, with every step versioned and publicly accessible. The experiment aims to test AI’s ability to manage complex, real-world business scenarios, emphasizing the importance of disciplined execution over mere diagnosis. The project also draws on prior research showing that AI models often recognize problems but struggle to complete tasks consistently in operational settings.

“Thorough analysis alone does not guarantee successful business outcomes; execution is the real challenge.”

— an anonymous researcher

Learning Robotic Process Automation: Create Software robots and automate business processes with the leading RPA tool – UiPath

Learning Robotic Process Automation: Create Software robots and automate business processes with the leading RPA tool – UiPath

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Operational Reliability

It remains unclear how these findings will translate to real-world businesses outside the experimental setup. The long-term impact of AI’s failure to complete actions, especially in high-pressure environments, is still being studied. Additionally, the experiment does not specify whether improvements in model training or interface design could mitigate these issues, leaving open the question of how to enhance AI’s execution capabilities effectively.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments in AI Management and Evaluation

The ongoing experiment will continue to track AI performance over multiple workdays, with updates on which models improve in completing actions. Researchers and practitioners will watch for strategies that bridge the gap between diagnosis and execution, potentially leading to new standards for AI operational reliability. The project may also influence how companies evaluate AI tools, emphasizing disciplined execution and verifiable results as key metrics.

Agentic AI Workflow Automation Engineering: Production-Ready Frameworks and Exercises for Platform Engineers

Agentic AI Workflow Automation Engineering: Production-Ready Frameworks and Exercises for Platform Engineers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of the Firmulate experiment?

The experiment aims to demonstrate the practical challenges of AI automation in managing a business, focusing on the gap between diagnosis and execution in real-time operations.

How does this experiment differ from typical AI demos?

Unlike standard demos that showcase isolated features, this project publicly runs a synthetic company, exposing ongoing decision-making, cash flow, and failures in real-time.

What are the key lessons from the latest results?

More thorough analysis does not necessarily lead to better outcomes; disciplined execution and evidence-based actions are crucial for AI success in real-world management.

Will this experiment influence how AI tools are used in business?

Yes, it highlights the importance of evaluating AI’s ability to complete actions reliably, which may shape future standards for operational AI deployment.

What remains uncertain about AI’s future in management?

It is still unclear how to develop AI systems that consistently turn diagnosis into successful, verifiable actions across diverse business contexts.

Source: ThorstenMeyerAI.com

You May Also Like
The Model Is Only 10%: The Real Lesson of the New SDLC

The Model Is Only 10%: The Real Lesson of the New SDLC

A new Google whitepaper reveals that the core of AI-driven software development isn’t the model itself, but how developers structure, verify, and control AI outputs.
The Real Cost Of A Local-Inference Rig In 2026

The Real Cost Of A Local-Inference Rig In 2026

Analyzing the true expenses of building a local AI inference setup in 2026, including hardware costs, VRAM constraints, and strategic choices for different model sizes.
The Real Cost of a Local-Inference Rig in 2026

The Real Cost of a Local-Inference Rig in 2026

A new analysis says local AI inference costs in 2026 hinge less on GPU speed than on VRAM capacity and value.
One Model, a Whole Portfolio: What Ten Days on Fable Mean for a Business Building on Frontier AI

One Model, a Whole Portfolio: What Ten Days on Fable Mean for a Business Building on Frontier AI

A developer ran nearly all his products through Anthropic’s Claude Fable 5 for ten days, revealing new AI-driven business capabilities and operational insights.