AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Diligence Is Not the Same as Delivery

We keep judging AI models by how impressively they think. A new public experiment suggests we should judge them by how reliably they finish — and the results are humbling for the very trait we praise most.

In Firmulate’s Crucible League, four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. One participant — Anthropic’s Opus 4.8 — was unambiguously the most thorough operator in the field. It wrote the deepest analyses, and it accumulated 80 learned playbook rules over the run, the most of any model.

It finished last.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Happened in the Crucible

The final league table, published after the July 2026 runs, reads: gpt-5.6-sol first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26; partial progress counts toward the score, but a single breach of trust caps the total, because in this experiment no amount of good work outweighs a breach of trust.

Here’s the remarkable part: all four models were excellent at the things we usually worry about. Every model spotted every crisis. Every model refused every manipulation attempt. If your fear is an AI agent going rogue or getting conned, this experiment offers reassurance — five out of five refusal performances, including a fake-CEO social engineering sequence that escalated over three stages and a reporter’s trap framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then came the €55,000 deal — and the gap. All the models diagnosed the same opportunity, made the same pitch, and only two of them signed it. Same diagnosis, same pitch, no signature from the others.

Amazon

CRM integration tools for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Split the Field

Why did only two close? The decisive competitive weakness — the fact that justified winning at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal. The models that didn’t, didn’t.

That’s a finding with legs beyond this experiment. If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well — Opus 4.8 writes beautifully. The question is whether they read your files first, and whether they finish what they start.

Amazon

AI business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: A Character Study in Effort Without Impact

To be clear and fair: Opus 4.8 was not a failure. It was the most diligent participant the experiment has measured — the deepest analyses, the most learned rules at 80, a genuine appetite for the work. What sunk it was two specific habits.

First, the close was left on the table. The deal its own analysis had earned went unsigned, because the decisive fact stayed buried in files it never dug through.

Second, discipline slipped under pressure: the model made write attempts into a locked department instead of escalating — trying to force a door rather than finding the person with the key. Thoroughness without prioritization; effort without judgment about where effort belongs.

The sting for the rest of the field: the same weakness appeared, weaker, in all four models. Opus 4.8 is not an outlier to gawk at — it’s the clearest expression of a tendency the whole generation shares.

One Fairness Footnote

The league also carries a disclosure worth knowing: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second with 93 points and what the results call the cleanest discipline of the field.

Amazon

AI deal closing support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch This Happen

What makes Firmulate different from a typical benchmark is that it’s alive. Thirteen synthetic employees operate a real-money-mechanics company — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned and auditable. You can watch the current runs at firmulate.com/benchmarks.html, and a quiz built from 242 real, unedited management decisions lets you try guessing which model made which call — a surprisingly effective way to feel the personality differences between them.

Enterprises can go further: a pilot program runs the same wargame against a read-only export of your own business, with nothing ever writing back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Lesson — for AI and for Us

The Opus 4.8 story is uncomfortable precisely because it’s familiar. Every organization has someone like this: the person with the longest memos, the most complete research, the most rules written — who still loses the deal because the close never happened and the locked door got pushed instead of escalated.

As we hand AI agents real operational work, the Crucible’s lesson is that volume of effort is not a proxy for impact. The winning models weren’t the ones that worked hardest; they were the ones that read the right document, asked for the signature, and escalated properly when blocked. Prioritization beats thoroughness — for AI, apparently, just as much as for people.

Before you hire an AI workforce, wargame it. The league table is public, the failures are instructive, and the most diligent candidate is not always the one that gets the job done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Did you feel it? Lake Michigan earthquake shakes Chicago area

A small earthquake was felt in the Chicago area near Lake Michigan, confirmed by residents and seismic data. Details are still emerging.

Earthquake Reported on Chicago’s North Shore, USGS Says Magnitude 2.9

USGS reports a magnitude 2.9 earthquake on Chicago’s North Shore, with details still emerging. Read for confirmed facts and next steps.

The Issus leafhopper is the only known creature in the natural world to have perfectly interlocking mechanical gears, which it uses to synchronize its legs for jumping.

Scientists confirm the Issus leafhopper is the only known creature with perfectly interlocking mechanical gears for synchronized jumping.

OlmoEarth Studio’s Embedding Exports: A Game Changer For AI Projects

OlmoEarth Studio now supports on-demand export of satellite data embeddings for improved land analysis, boosting AI applications in earth observation.