AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A chatbot can sound decisive in a demo. But what happens when an AI has to run a business through a week of churn, pressure and tempting shortcuts? Firmulate’s experiment puts models in that position—and finds a gap between knowing what to do and doing it.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, crises and temptations. Their decisions were versioned and auditable. The experiment was designed to reveal management quality under pressure, not just polished chat.

The league ended with gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Good diagnosis, unfinished business

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s succinct verdict: “Same diagnosis, same pitch — no signature.”

The missed opportunity came down to a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a revealing kind of failure: the right analysis can be present, while the decisive action still goes undone.

Trust held up under a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

What the leaderboard leaves out

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is also a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results make for a useful comparison, but that difference belongs alongside the rankings.

Firmulate presents the experiment as a live company, not a fictional case study. Its synthetic workforce has 13 employees, real money mechanics, a burn of €105k a month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.

From watching to trying it on your business

The enterprise proposition is to move from observing a live experiment to running a wargame against a company’s own circumstances. A pilot uses a read-only export to create a digital twin, then tests crisis scenarios and reports model rankings alongside weak points in the company’s playbooks. Nothing writes back to real systems.

That boundary matters: a business can examine how models respond to its customers, pipeline, rules and imagined crises without giving the experiment permission to alter operational tools. The results can help leaders see not only whether a model recognizes a problem, but whether it follows through, escalates appropriately and respects trust under pressure.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Firmulate’s league suggests that spotting a crisis and refusing a scam are only part of the job. Finding the evidence, closing the deal and following the rules also matter. Enterprises can run the wargame against a read-only export of their own business, with no writes to live systems. Explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Issus leafhopper is the only known creature in the natural world to have perfectly interlocking mechanical gears, which it uses to synchronize its legs for jumping.

Scientists confirm the Issus leafhopper is the only known creature with perfectly interlocking mechanical gears for synchronized jumping.

How to See the Giant Asteroid That Will Pass by Earth This Weekend

Learn how to see asteroid 1997 NC1 as it passes 2.56 million km from Earth on June 27, with viewing tips and timing details.

Exploring the island where nearly half of Japan’s lead is produced

Chigirishima Island in Japan’s Seto Inland Sea produces nearly half of the nation’s lead, highlighting its industrial significance and ongoing operations.

Weathergotchi – An E-Paper Climate Logger

A new eco-friendly device called Weathergotchi, an E-Paper climate logger, has been introduced to monitor local weather and environmental data.