
The A-Student Who Never Signs the Contract
We’ve all met one: the brilliant colleague who diagnoses every problem perfectly, drafts the flawless proposal, and then… never asks for the signature. According to a live experiment by Firmulate, today’s frontier AI models are exactly that colleague.
Four frontier models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed, and every decision was versioned and auditable. The results, finalized in July 2026 in what Firmulate calls its Crucible League, expose a measurement gap that coding leaderboards and chat arenas completely miss.
AI customer relationship management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Scoreboard
The final standings: gpt-5.6-sol took first with 95 points, Kimi K3 — the newcomer from Moonshot — scored 93, Sonnet 5 landed at 88, Fable 5 at 77, and Opus 4.8 finished last at 73. For context, a do-nothing baseline still scores 26, because partial progress counts. But there’s a hard ceiling built into the scoring philosophy: a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”
One fairness note worth flagging: K3 ran without an effort parameter (API default) while its rivals ran at xhigh — and still nearly won.
AI deal-closing automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here’s the finding that should stop every executive scroll. All four models spotted every crisis. All four refused every manipulation attempt thrown at them. And yet only two of the five managed to sign the €55,000 deal that their own analysis had earned. Firmulate’s terse summary: “Same diagnosis, same pitch — no signature.”
That gap is invisible in chat demos. A model can be eloquent, technically sharp, and politically correct in every answer — and still leave the close on the table when it matters.
The Buried Fact
The deal-turning detail wasn’t in any customer conversation. It sat two document references deep in the company’s own files: a decisive competitor weakness that the customer never mentioned. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson maps directly onto human management: the answer is often already in your own archives, if anyone bothers to look.
As an affiliate, we earn on qualifying purchases.
They All Passed the Integrity Test
The social engineering gauntlet deserves its own applause. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Cautionary Tale: Opus 4.8
The most striking individual profile belongs to the last-place finisher. Opus 4.8 was the most thorough participant of the entire field — it learned over 80 new rules and produced the deepest analyses. Yet it finished last: the close was left unclaimed, and discipline slipped, including write attempts into a locked department instead of escalating the issue. And here’s the uncomfortable part — the same weakness appeared, weaker, in all four models. Thoroughness and follow-through are not the same muscle.
As an affiliate, we earn on qualifying purchases.
Not a Slide Deck — a Company You Can Watch
Firmulate isn’t running this on static test data. The live company has 13 synthetic employees and real money mechanics: it burns €105k per month against just €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com/live.
Want to test your own instincts? A quiz built from 242 real, unedited management decisions lets you guess which model made which call. Full benchmarks and plain-language findings are at firmulate.com/benchmarks.html. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

Management Quality, Not Chat Quality
Coding benchmarks measure whether an AI can solve a problem. Chat arenas measure whether it answers nicely. Neither measures what happens when an agent touches your CRM, your support queue, or your forecast under real pressure: does it finish what it starts, does it read your files first, does it stay honest when a fake CEO comes knocking — and what does a unit of useful work actually cost?
The Crucible League suggests we’ve been grading AI on the easy part. Spotting crises and refusing manipulation turned out to be table stakes for every frontier model. Closing the deal — the boring, procedural, finish-the-job part — separated the winners from the also-rans. Before you hand an AI agent real responsibility, ask the question Firmulate is built around: not “does it write well,” but “can it manage?”
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html