
A chatbot can sound decisive in a demo. But what happens when an AI has to run a business through a week of churn, pressure and tempting shortcuts? Firmulate’s experiment puts models in that position—and finds a gap between knowing what to do and doing it.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, crises and temptations. Their decisions were versioned and auditable. The experiment was designed to reveal management quality under pressure, not just polished chat.
The league ended with gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Good diagnosis, unfinished business
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s succinct verdict: “Same diagnosis, same pitch — no signature.”
The missed opportunity came down to a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a revealing kind of failure: the right analysis can be present, while the decisive action still goes undone.
Trust held up under a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
What the leaderboard leaves out
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is also a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results make for a useful comparison, but that difference belongs alongside the rankings.
Firmulate presents the experiment as a live company, not a fictional case study. Its synthetic workforce has 13 employees, real money mechanics, a burn of €105k a month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.
From watching to trying it on your business
The enterprise proposition is to move from observing a live experiment to running a wargame against a company’s own circumstances. A pilot uses a read-only export to create a digital twin, then tests crisis scenarios and reports model rankings alongside weak points in the company’s playbooks. Nothing writes back to real systems.
That boundary matters: a business can examine how models respond to its customers, pipeline, rules and imagined crises without giving the experiment permission to alter operational tools. The results can help leaders see not only whether a model recognizes a problem, but whether it follows through, escalates appropriately and respects trust under pressure.

Put your playbooks to the test
Firmulate’s league suggests that spotting a crisis and refusing a scam are only part of the job. Finding the evidence, closing the deal and following the rules also matter. Enterprises can run the wargame against a read-only export of their own business, with no writes to live systems. Explore the Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
