firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the fast-evolving world of artificial intelligence, one surprisingly revealing fact emerges: even a completely inactive AI baseline scores 26 out of 100 in a rigorous business benchmark. For managers and decision-makers, understanding this baseline is crucial. It highlights how much AI models are expected to do — and how much they actually accomplish under real-world pressures.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Chat Quality

At first glance, AI performance might seem like a matter of how well the model can generate convincing conversation or support creative tasks. But the latest public experiment from Firmulate reveals a deeper truth: managing AI in a business context means ensuring it can handle crises, resist manipulation, and follow complex protocols — not just produce pretty words.

The Unseen Power of Partial Progress

The benchmark scores all models on a scale up to 100, but what’s truly striking is that a do-nothing baseline, which refuses to act or make decisions, scores 26 points. This is because even minimal engagement — such as reading files or refusing manipulative requests — counts towards the score. It’s a reminder that stopping short isn’t enough; models must actively demonstrate integrity and problem-solving ability.

Why a Single Breach Caps Performance

The experiment makes it clear: if an AI model breaches trust — say, by signing a fraudulent deal or bypassing security protocols — its maximum achievable score is capped. This safeguard ensures genuine diligence and honesty are rewarded, rather than superficial compliance or careless mistakes.

Amazon

AI model performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Wargame: Real Business, Real Crises

In a real-time test conducted in a simulated company environment, four leading AI models were tasked with running a small software firm facing the toughest week imaginable. They had to navigate customer crises, internal risks, manipulative tactics, and complex decision-making — all with transparent, auditable processes.

Consistent Crisis Detection and Rejection of Manipulation

Remarkably, all models identified every crisis and refused every manipulation attempt, including elaborate social engineering schemes like staged CEO messages and reporter tricks. Kimi K3, the most disciplined in the group, explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Hidden Weakness: Reading Files Matters

The key difference in performance stemmed from how deeply models could read and interpret internal documents. The most successful model, gpt-5.6-sol, found a buried reference in a file that led to a €4,583 million deal at full price. Others missed this crucial detail because they either didn’t read deep enough or failed to escalate properly.

Discipline and Process Slips

Even the most thorough participant, Opus 4.8, left the deal on the table by slipping into a locked department instead of escalating issues — a sign that discipline and process adherence remain vital. Interestingly, models without effort parameters (API defaults) performed slightly less disciplined, but the core findings remained consistent across configurations.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Deployment

For companies considering AI assistants or automation, the takeaway is clear: it’s not just about how well they chat but whether they can finish tasks, stay honest, and read internal files properly. A model that can’t detect a simple manipulation or read a crucial file risks making costly errors or being exploited.

Measuring What Matters

Firmulate’s live benchmark openly displays ongoing results, with models continuously tested against real crises and temptations. The league table shows gpt-5.6-sol at 95 points, Kimi K3 at 93, and others trailing behind. The scores reflect actual performance, not just language fluency, emphasizing trustworthiness and thoroughness.

Accessible Testing for Businesses

Enterprises can run their own wargames using a read-only export of their systems, ensuring they can evaluate AI performance without risking real data. The public platform makes this testing transparent and accessible, encouraging responsible AI adoption.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Rep. Nancy Pelosi’s husband, Paul, faces charge for Napa County hit-and-run

Paul Pelosi, husband of Rep. Nancy Pelosi, faces charges after a hit-and-run incident in Napa County, California, according to authorities.

Lanark, South Lanarkshire, United Kingdom Surges In Global Coverage

Lanark, South Lanarkshire, UK, experiences a surge in international coverage, with 16 mentions in recent media monitoring, highlighting increased global interest.

Severe Thunderstorm Warning Issued August 29 At 11:54PM CDT Until August 30 At 12:00AM CDT By NWS Aberdeen SD

A severe thunderstorm warning was issued for Dewey, SD, on August 29 at 11:54 PM CDT, lasting until August 30 at 1:00 AM, according to NWS.

Russia’s Ambassador To Germany Says A Leipzig Airport Attack Would Have Looked ‘Completely Different’ If Moscow Had Wanted To Carry One Out

Russia’s ambassador to Germany states a Leipzig airport attack would have looked ‘completely different’ if Moscow had planned it, highlighting geopolitical tensions.