firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

As AI becomes a more integral part of business operations, the real question isn’t just whether these systems can handle a task—it’s whether they can see it through. A recent public experiment with four top AI models reveals surprising insights into how diligence and prioritization impact actual business performance.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Simulated Business Crisis

In a groundbreaking live experiment, four leading AI models were tasked with managing a small software company through its worst week—full of crises, customer demands, and potential manipulation attempts. The goal? See if the AI could identify critical issues, resist dishonest tactics, and ultimately close a lucrative deal worth €55,000.

Every decision made by these AI agents was meticulously versioned and auditable, providing transparency into their reasoning processes. The models included the well-known GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, with scores ranging from 73 to 95 in the Crucible League benchmarks—scores that reflect their overall capability in complex decision-making scenarios.

The Core Findings: Diligence Doesn’t Guarantee Impact

All four models demonstrated impressive technical prowess—they identified every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. In fact, none of them fell prey to these common workplace traps, showcasing integrity at a fundamental level.

However, only two models actually secured the deal at full price, despite their keen crisis detection. The models that succeeded—GPT-5.6-sol and Kimi K3—did so by thoroughly analyzing the company’s files to uncover a buried but decisive piece of information, located two document references deep. This insight was crucial; reading the files won the deal, worth an additional €4,583 of monthly recurring revenue (MRR). In contrast, the less thorough models, including Sonnet 5 and Fable 5, failed to uncover this. Their discipline slipped, and they left the critical opportunity on the table.

The Hidden Weakness: Not All Diligence Is Equal

Interestingly, the most thorough participant—Opus 4.8—had over 80 learned rules and performed the deepest analyses. Yet it finished last, primarily because its discipline slipped during the close phase, recording write attempts into a locked department instead of escalating them. This underlines a vital point: volume of rules and depth of analysis are not enough. Prioritization and discipline in execution matter just as much, if not more.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and AI Adoption

For businesses contemplating integrating AI into critical decision-making roles, these results are illuminating. The models are capable of recognizing crises and resisting manipulation—an essential baseline. But the true test lies in their ability to prioritize, analyze deeply, and stay disciplined when it counts most.

In the live experiment, the models’ scores ranged from 73 to 95. The top performer, GPT-5.6-sol, achieved 95, while Kimi K3 scored 93 and was praised for its clean discipline. Meanwhile, models like Sonnet 5 and Fable 5 scored lower, 88 and 77 respectively, with process slips contributing to their failure to close the deal.

Social Engineering Resistance: A Clear Success

All models refused social engineering attempts—such as escalating fake CEO messages or background-only approval requests—a critical feature for enterprise security. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating a cautious, rule-based approach that prioritized security over compliance.

Amazon

business AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Bigger Picture: Diligence Isn’t Enough

The experiment underscores that diligence—number of rules learned or depth of analysis—is not enough to guarantee success. Impact depends heavily on prioritization, discipline, and the ability to focus on what truly matters. The instance of the most thorough model missing the deal because it failed to escalate and escalate appropriately exemplifies this core lesson.

As AI continues to collaborate with humans in high-stakes environments, understanding how to align diligence with impact becomes vital. The experiment suggests that models must be trained not just to analyze exhaustively, but to prioritize effectively and maintain discipline under pressure.

Amazon

enterprise AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Looking Ahead: Wargaming Your AI Workforce

For enterprises eager to test their AI systems before deployment, the Firmulate platform offers a unique opportunity. By running the same business scenario against their AI, companies can observe how the models perform in real crises—without risking their actual operations. This “wargaming” enables organizations to identify whether their AI agents can stay honest, focus on vital issues, and deliver real business results.

In conclusion, the experiment paints a clear picture: in AI-enabled decision-making, volume and depth matter less than focus, discipline, and prioritization. Achieving impactful results requires more than thoroughness; it demands strategic discipline and trustworthiness, especially when the stakes are high.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Afghanistan Surges In Global Coverage

Recent data shows Afghanistan’s coverage has surged in international media, marking a notable shift in global attention. Details and implications explained.

Cdu Friedrich Merz

Friedrich Merz has secured a narrow victory in the CDU leadership election, reaffirming his position amid internal party debates and upcoming elections.

Trump Seeks to Undo USMCA, but Breaking Deal Could Cost Billions

Former President Trump proposes to reverse the USMCA agreement, but experts warn that withdrawing could lead to significant economic penalties for the U.S.

Tornado Warning Issued August 9 At 11:45PM CDT Until August 10 At 12:00AM CDT By NWS Chicago IL

A tornado warning was issued for La Salle, Illinois, on August 9 at 11:45 PM CDT, effective until 12:00 AM CDT on August 10, according to NWS Chicago.