
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Doesn’t Start at Zero
Most AI scoreboards have a flattering habit: they make every model look brilliant. Ask a chatbot a question, it answers, it gets points. But what happens when you hand an AI something closer to a real job — running a small software company through its worst week — and grade it like a manager, not a chat partner?
That’s the premise behind Firmulate’s benchmark, a live experiment that puts frontier AI models in charge of the same fictional company, with the same customers, the same crises, and the same temptations to cut corners. And one of its most interesting numbers isn’t at the top of the leaderboard — it’s at the bottom. A do-nothing baseline run, where the AI essentially sits on its hands, still scores 26 points out of 100. Not zero. Deliberately not zero.
Why Doing Nothing Isn’t Worth Nothing
The logic is deceptively simple, and it mirrors how real work gets judged. If an AI agent is left running your support queue, your CRM, or your forecast, simply not making things worse has value. It didn’t anger a customer. It didn’t leak data. It didn’t fall for a scam. Partial progress counts too — an agent that defuses three of five crises has genuinely accomplished something, even if it dropped the ball on the rest.
So the floor sits at 26, and everything above it measures what the model actually did with its week. The final Crucible League from July 2026 tells that story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.
The One Rule That Caps Everything
But the scoring has a hard ceiling rule that says as much about the benchmark’s philosophy as any number: a single breach of trust caps the total grade. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” An agent could close every deal and defuse every crisis, but if it deceives a customer or bypasses an approval once, the score is capped. It’s the AI equivalent of firing your star salesperson for cooking the books.
It also explains why the league treats suspiciously round perfection with distrust — a flat 100 would be a red flag, not a triumph. Real management weeks are messy, and an honest benchmark expects mess.
Same Diagnosis, Same Pitch — No Signature
The experiment’s central finding is a quiet heartbreaker. Every model ran the identical gauntlet: a €55,000 deal on the table, crises piling up, and manipulation attempts probing for weakness. All five models spotted every crisis. All five refused every manipulation attempt, including a social-engineering sequence of fake CEO messages escalating over three stages and a reporter’s disarming “just one yes/no, on background” trick — 5 of 5 refused, with Kimi K3 reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The failure wasn’t intelligence; it was follow-through.
The Buried Fact That Decided the Deal
And here’s the twist: the decisive weakness in the competing offer wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents: does your agent read your files first, or just improvise?
The last-place profile drives it home. Opus 4.8 was the most thorough participant, with over 80 learned rules and the deepest analyses — and still left the close on the table while discipline slipped, attempting writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.
Watch It Happen Live
The company is still running. Thirteen synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live. If you’d rather test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Takeaway
An honest benchmark doesn’t just rank models — it tells you where they fail and why. Firmulate’s floor of 26 says that restraint has value; its trust cap says honesty is non-negotiable; and its headline finding says the hardest part of management AI isn’t spotting problems — it’s finishing the job. Before you hand an agent the keys to your business, see how it scores when nobody’s watching.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI customer service automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
