
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Diligence Is Not the Same as Delivery
We keep judging AI models by how impressively they think. A new public experiment suggests we should judge them by how reliably they finish — and the results are humbling for the very trait we praise most.
In Firmulate’s Crucible League, four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. One participant — Anthropic’s Opus 4.8 — was unambiguously the most thorough operator in the field. It wrote the deepest analyses, and it accumulated 80 learned playbook rules over the run, the most of any model.
It finished last.
As an affiliate, we earn on qualifying purchases.
What Actually Happened in the Crucible
The final league table, published after the July 2026 runs, reads: gpt-5.6-sol first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26; partial progress counts toward the score, but a single breach of trust caps the total, because in this experiment no amount of good work outweighs a breach of trust.
Here’s the remarkable part: all four models were excellent at the things we usually worry about. Every model spotted every crisis. Every model refused every manipulation attempt. If your fear is an AI agent going rogue or getting conned, this experiment offers reassurance — five out of five refusal performances, including a fake-CEO social engineering sequence that escalated over three stages and a reporter’s trap framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then came the €55,000 deal — and the gap. All the models diagnosed the same opportunity, made the same pitch, and only two of them signed it. Same diagnosis, same pitch, no signature from the others.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Split the Field
Why did only two close? The decisive competitive weakness — the fact that justified winning at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal. The models that didn’t, didn’t.
That’s a finding with legs beyond this experiment. If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well — Opus 4.8 writes beautifully. The question is whether they read your files first, and whether they finish what they start.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: A Character Study in Effort Without Impact
To be clear and fair: Opus 4.8 was not a failure. It was the most diligent participant the experiment has measured — the deepest analyses, the most learned rules at 80, a genuine appetite for the work. What sunk it was two specific habits.
First, the close was left on the table. The deal its own analysis had earned went unsigned, because the decisive fact stayed buried in files it never dug through.
Second, discipline slipped under pressure: the model made write attempts into a locked department instead of escalating — trying to force a door rather than finding the person with the key. Thoroughness without prioritization; effort without judgment about where effort belongs.
The sting for the rest of the field: the same weakness appeared, weaker, in all four models. Opus 4.8 is not an outlier to gawk at — it’s the clearest expression of a tendency the whole generation shares.
One Fairness Footnote
The league also carries a disclosure worth knowing: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second with 93 points and what the results call the cleanest discipline of the field.
As an affiliate, we earn on qualifying purchases.
You Can Watch This Happen
What makes Firmulate different from a typical benchmark is that it’s alive. Thirteen synthetic employees operate a real-money-mechanics company — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned and auditable. You can watch the current runs at firmulate.com/benchmarks.html, and a quiz built from 242 real, unedited management decisions lets you try guessing which model made which call — a surprisingly effective way to feel the personality differences between them.
Enterprises can go further: a pilot program runs the same wargame against a read-only export of your own business, with nothing ever writing back to real systems.

The Lesson — for AI and for Us
The Opus 4.8 story is uncomfortable precisely because it’s familiar. Every organization has someone like this: the person with the longest memos, the most complete research, the most rules written — who still loses the deal because the close never happened and the locked door got pushed instead of escalated.
As we hand AI agents real operational work, the Crucible’s lesson is that volume of effort is not a proxy for impact. The winning models weren’t the ones that worked hardest; they were the ones that read the right document, asked for the signature, and escalated properly when blocked. Prioritization beats thoroughness — for AI, apparently, just as much as for people.
Before you hire an AI workforce, wargame it. The league table is public, the failures are instructive, and the most diligent candidate is not always the one that gets the job done.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.