
In the fast-evolving world of artificial intelligence, one surprisingly revealing fact emerges: even a completely inactive AI baseline scores 26 out of 100 in a rigorous business benchmark. For managers and decision-makers, understanding this baseline is crucial. It highlights how much AI models are expected to do — and how much they actually accomplish under real-world pressures.
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: More Than Just Chat Quality
At first glance, AI performance might seem like a matter of how well the model can generate convincing conversation or support creative tasks. But the latest public experiment from Firmulate reveals a deeper truth: managing AI in a business context means ensuring it can handle crises, resist manipulation, and follow complex protocols — not just produce pretty words.
The Unseen Power of Partial Progress
The benchmark scores all models on a scale up to 100, but what’s truly striking is that a do-nothing baseline, which refuses to act or make decisions, scores 26 points. This is because even minimal engagement — such as reading files or refusing manipulative requests — counts towards the score. It’s a reminder that stopping short isn’t enough; models must actively demonstrate integrity and problem-solving ability.
Why a Single Breach Caps Performance
The experiment makes it clear: if an AI model breaches trust — say, by signing a fraudulent deal or bypassing security protocols — its maximum achievable score is capped. This safeguard ensures genuine diligence and honesty are rewarded, rather than superficial compliance or careless mistakes.
AI model performance evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Wargame: Real Business, Real Crises
In a real-time test conducted in a simulated company environment, four leading AI models were tasked with running a small software firm facing the toughest week imaginable. They had to navigate customer crises, internal risks, manipulative tactics, and complex decision-making — all with transparent, auditable processes.
Consistent Crisis Detection and Rejection of Manipulation
Remarkably, all models identified every crisis and refused every manipulation attempt, including elaborate social engineering schemes like staged CEO messages and reporter tricks. Kimi K3, the most disciplined in the group, explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Hidden Weakness: Reading Files Matters
The key difference in performance stemmed from how deeply models could read and interpret internal documents. The most successful model, gpt-5.6-sol, found a buried reference in a file that led to a €4,583 million deal at full price. Others missed this crucial detail because they either didn’t read deep enough or failed to escalate properly.
Discipline and Process Slips
Even the most thorough participant, Opus 4.8, left the deal on the table by slipping into a locked department instead of escalating issues — a sign that discipline and process adherence remain vital. Interestingly, models without effort parameters (API defaults) performed slightly less disciplined, but the core findings remained consistent across configurations.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Deployment
For companies considering AI assistants or automation, the takeaway is clear: it’s not just about how well they chat but whether they can finish tasks, stay honest, and read internal files properly. A model that can’t detect a simple manipulation or read a crucial file risks making costly errors or being exploited.
Measuring What Matters
Firmulate’s live benchmark openly displays ongoing results, with models continuously tested against real crises and temptations. The league table shows gpt-5.6-sol at 95 points, Kimi K3 at 93, and others trailing behind. The scores reflect actual performance, not just language fluency, emphasizing trustworthiness and thoroughness.
Accessible Testing for Businesses
Enterprises can run their own wargames using a read-only export of their systems, ensuring they can evaluate AI performance without risking real data. The public platform makes this testing transparent and accessible, encouraging responsible AI adoption.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
