The AI Company Turning Corporate Survival Into A Live Feed
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Firmulate has launched a live experiment with a synthetic workforce managing a software company. The project reveals that while AI can diagnose problems, completing actions remains a critical challenge. This experiment offers new insights into AI’s practical capabilities and limitations, as detailed in the original analysis.

Firmulate has launched a live experiment where a synthetic workforce of 13 AI agents manages a software company, exposing the practical challenges of automation in real-time. This public project reveals that, despite sophisticated analysis and rules, AI systems often fail to complete critical business actions, risking company survival. The experiment’s transparency offers a rare view into AI’s operational limits and the importance of disciplined execution.

The experiment involves a synthetic team operating a company with a monthly burn rate of €105,000 against €2,300 in recurring revenue, illustrating the challenges of turning nonprofits into companies. Every workday is versioned, creating an evolving record of decisions, actions, successes, and failures. Over 680 self-learned rules have been developed, yet the project demonstrates that thorough analysis alone does not guarantee successful business outcomes. The core issue remains the AI’s ability to turn diagnosis into action, with many models recognizing problems but failing to complete critical tasks. This aligns with the insights from the original analysis.

In the latest results, only two of five AI models secured a €55,000 deal, despite all identifying the opportunity. The decisive factor was uncovering a hidden weakness in the customer’s decision process, buried in the company’s files—an insight that only some models followed through to act upon. The experiment also tested trust, with AI agents refusing fake approval requests, showing that trust is maintained through evidence retrieval and disciplined action rather than superficial compliance. The top-performing model, gpt-5.6-sol, scored 95 out of 100, while the most thorough, Opus 4.8, scored only 73, highlighting that more analysis does not necessarily lead to better results.

At a glance
breakingWhen: ongoing, with real-time updates availab…
The developmentFirmulate is running a live AI-driven company with 13 synthetic employees, exposing the gap between AI diagnosis and execution amid ongoing cash pressure.

Implications of AI’s Inability to Complete Business Actions

This experiment underscores a critical challenge for AI automation: diagnosis and recognition of problems are insufficient without reliable execution. For businesses, it highlights that AI’s value depends on its capacity to follow through on decisions, resist pressure, and respect organizational boundaries. The live, public nature of the experiment makes these limitations visible, offering a new lens for evaluating AI tools beyond their analytical capabilities. It also raises questions about the future of AI-driven management and the importance of disciplined, verifiable actions for organizational survival.

Amazon

AI automation management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Live Automation Experiments

Firmulate’s project builds on the trend of transparency in AI development, where companies share their operational experiments publicly. Unlike typical demos or product launches, this ongoing experiment involves a synthetic workforce managing a real company’s decisions and cash flow, with every step versioned and publicly accessible. The experiment aims to test AI’s ability to manage complex, real-world business scenarios, emphasizing the importance of disciplined execution over mere diagnosis. The project also draws on prior research showing that AI models often recognize problems but struggle to complete tasks consistently in operational settings.

“Thorough analysis alone does not guarantee successful business outcomes; execution is the real challenge.”

— an anonymous researcher

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Operational Reliability

It remains unclear how these findings will translate to real-world businesses outside the experimental setup. The long-term impact of AI’s failure to complete actions, especially in high-pressure environments, is still being studied. Additionally, the experiment does not specify whether improvements in model training or interface design could mitigate these issues, leaving open the question of how to enhance AI’s execution capabilities effectively.

Amazon

business process automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments in AI Management and Evaluation

The ongoing experiment will continue to track AI performance over multiple workdays, with updates on which models improve in completing actions. Researchers and practitioners will watch for strategies that bridge the gap between diagnosis and execution, potentially leading to new standards for AI operational reliability. The project may also influence how companies evaluate AI tools, emphasizing disciplined execution and verifiable results as key metrics.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of the Firmulate experiment?

The experiment aims to demonstrate the practical challenges of AI automation in managing a business, focusing on the gap between diagnosis and execution in real-time operations.

How does this experiment differ from typical AI demos?

Unlike standard demos that showcase isolated features, this project publicly runs a synthetic company, exposing ongoing decision-making, cash flow, and failures in real-time.

What are the key lessons from the latest results?

More thorough analysis does not necessarily lead to better outcomes; disciplined execution and evidence-based actions are crucial for AI success in real-world management.

Will this experiment influence how AI tools are used in business?

Yes, it highlights the importance of evaluating AI’s ability to complete actions reliably, which may shape future standards for operational AI deployment.

What remains uncertain about AI’s future in management?

It is still unclear how to develop AI systems that consistently turn diagnosis into successful, verifiable actions across diverse business contexts.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
15 Best AI-Powered Student Planners For Academic Organization In 2026

15 Best AI-Powered Student Planners For Academic Organization In 2026

Discover the 15 best AI-driven student planners for academic success in 2026, including features, benefits, and what to consider before choosing.
A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

Anthropic reveals that effective AI Skills are structured as folders containing instructions, scripts, and assets, transforming organizational workflows.
The Weights Came First: What Thinking Machines’ Inkling Actually Signals

The Weights Came First: What Thinking Machines’ Inkling Actually Signals

Thinking Machines released the full weights of Inkling under Apache 2.0, making it the first foundation model from the lab available openly before competing models.
The Free-Download Question: When Running Your Own Model Actually Beats Paying

The Free-Download Question: When Running Your Own Model Actually Beats Paying

Analysis of when owning and operating open-weight AI models is more cost-effective than paying API fees, considering hardware, operational costs, and performance.