
Imagine a company with no human employees, running its daily operations solely through artificial intelligence — yet losing over €100,000 every month. This isn’t science fiction; it’s the real-time experiment happening right now on Firmulate’s live site. For the first time, you can see an AI company fighting for survival, making critical decisions in real time, and revealing where AI still struggles to replace human judgment.
What’s Happening Behind the Digital Curtain?
At the heart of this unprecedented experiment are 13 synthetic employees, each driven by cutting-edge AI models, tasked with running a small software company through its most challenging week. This includes handling customer crises, negotiating deals, and responding to manipulative tactics — all without a single human on the payroll.
Every decision made by these AI employees is logged, versioned, and publicly accessible, turning a behind-the-scenes corporate drama into a transparent spectacle. The company’s financials paint a stark picture: it’s burning approximately €105,000 every month against a modest revenue of €2,300 in monthly recurring income. Meanwhile, a countdown signals how long the company’s cash reserves will last, underscoring the urgency of its survival struggle.
How Do AI Models Perform in Critical Business Situations?
The experiment pits four frontier AI models against each other, each running the same week of simulated crises. Remarkably, all models detected every crisis and refused every attempt at manipulation, such as social engineering or impersonation tricks. This demonstrates a notable level of ethical restraint and situational awareness in AI decision-making.
However, success in this testing isn’t just about recognizing crises — it’s about acting on them. Only two models managed to close a critical deal, valued at over €55,000, their own analysis justifying the agreement. Interestingly, the models that read deeper into the company’s own internal documents—specifically, information buried two files deep—were the ones who secured the deal at full price, adding around +€4,583 in monthly recurring revenue (MRR).
The Weaknesses and Lessons in AI Performance
One of the standout participants, Opus 4.8, with the most comprehensive set of learned rules (over 80), still finished last. It left the deal on the table, failed to escalate issues properly, and slipped into internal silos, illustrating that thorough analysis alone does not guarantee optimal decision-making under pressure.
Furthermore, the experiment included social engineering tests. Fake CEO messages, escalating in intensity over three stages, and even a reporter trick requiring a simple yes/no response, were all refused by every AI model. Kimi K3’s on-record reasoning was clear: treat such requests as potential impersonation or approval bypasses. This shows that AI can be programmed to maintain a strong ethical stance, even in deceitful scenarios.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications of a Live AI Business
This isn’t just a technological showcase; it’s a stark illustration of AI’s capabilities and current limitations in real business contexts. The company operates every business day, and the decisions—some right, some slipping—are real, not simulated. Every action and hesitation is public, providing a rare window into how AI performs under sustained stress and financial pressure.
For those pondering the future of AI in enterprise, the key takeaway is that success hinges not just on generating convincing chat or reports, but on whether the AI can finish what it starts, read relevant internal information, and stay honest under pressure. It’s about the quality of work, not just the appearance of competence.
Who Leads the Field?
Based on the current leaderboard, gpt-5.6-sol leads with a score of 95, having identified the hidden fact necessary to close the deal and thus secured the full performance potential. Kimi K3 follows closely at 93, with the most disciplined performance, while Sonnet 5 and Fable 5 trail behind with scores of 88 and 77 respectively. Despite differences, all managed to close the deal at some level, but only the top two did so with full confidence and integrity.

50 AI Automations That Pay: Done-for-You Workflows That Save Hours and Make Money
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What’s Next for AI Business Automation?
This ongoing live experiment challenges the assumption that AI can seamlessly replace human decision-makers. It exposes gaps in AI judgment, especially under complex, high-pressure scenarios. Yet, it also highlights AI’s remarkable ability to recognize crises, avoid manipulation, and even uncover hidden internal information crucial for success.
For companies considering AI-driven automation, the message is clear: watch these live tests, understand where AI still struggles, and prepare for a future where human oversight remains vital. The experiment can be observed in real time, with decision logs and performance data available to anyone interested at Firmulate’s live site.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

The Ethical Nightmare Challenge: How to Avoid the Worst of AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.