firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In a newsroom, the facts can be right and the story can still fall short

A reporter may find the key document, verify the central claim and still miss the detail that turns a report into a scoop. Businesses face a similar test as they hand more work to AI: can a model not only identify a crisis, but make the decision that resolves it? Firmulate’s live company experiment puts that question on display, with a path from watching AI at work to testing it against a company’s own risks.

A company under pressure, in public

Firmulate runs AI models as complete companies, with synthetic employees handling workdays, crises and money mechanics. Its live company has 13 synthetic employees. It faces a monthly burn of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the experiment can be watched at firmulate.com.

The company’s final Crucible League, in July 2026, put frontier models through the same worst week at a small software company: the same customers, crises and temptations. The results ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s rule is blunt: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary captures the miss: “Same diagnosis, same pitch — no signature.” For a business, that gap matters. A model can produce the right assessment and still leave a consequential decision unfinished.

The deal also depended on a clue hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding turns attention from fluent explanations to whether an AI can find and act on relevant information in the material a company actually holds.

Trust under pressure

The experiment tested social engineering with fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. The comparison suggests that diligence alone does not guarantee sound execution.

There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also makes 242 real, unedited management decisions available in a “guess the model” quiz at firmulate.com, giving readers a way to inspect the choices behind the rankings.

From watching to a company’s own test

A public experiment can show how models behave in one company. A pilot can ask what happens when the stakes and playbooks belong to yours. Enterprises can run the wargame against a read-only export of their business, with crisis scenarios drawn against that context and a board report covering model rankings and weaknesses in the company’s own playbooks. Nothing writes back to real systems.

That makes the proposal a practical next step for organizations weighing AI agents for customer records, support or forecasting: test decisions against company-specific conditions before relying on them in live operations. The live company is watchable; the pilot is a chance to see how models handle your company’s particular pressures.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s experiment shows that spotting a crisis and refusing manipulation are only part of the job; models can still miss the decision that closes the deal. To run a pilot against a read-only export of your business and receive a board report on model performance and playbook weaknesses, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ukrainians Surges In Global Coverage

Analysis of recent data shows a significant increase in global media coverage of Ukraine, with potential implications for international awareness and policy.

Elizondo Confirms Gubernatorial Ambitions Amid Morena Tensions

Luis Elizondo announces his bid for governor of Nuevo León, amid rising tensions with Morena, marking a significant political shift in the state.

Broad Peak, Pakistan (General), Pakistan Surges In Global Coverage

Pakistan’s Broad Peak experiences a surge in international media coverage, with 32 mentions in recent reports, highlighting increased global interest.

Live updates: Keir Starmer announces he’ll resign as UK prime minister, launching contest for successor

Keir Starmer has announced he will step down as UK Prime Minister and initiate a leadership race, sparking political upheaval amid ongoing challenges.