AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Start at Zero

Most AI scoreboards have a flattering habit: they make every model look brilliant. Ask a chatbot a question, it answers, it gets points. But what happens when you hand an AI something closer to a real job — running a small software company through its worst week — and grade it like a manager, not a chat partner?

That’s the premise behind Firmulate’s benchmark, a live experiment that puts frontier AI models in charge of the same fictional company, with the same customers, the same crises, and the same temptations to cut corners. And one of its most interesting numbers isn’t at the top of the leaderboard — it’s at the bottom. A do-nothing baseline run, where the AI essentially sits on its hands, still scores 26 points out of 100. Not zero. Deliberately not zero.

Why Doing Nothing Isn’t Worth Nothing

The logic is deceptively simple, and it mirrors how real work gets judged. If an AI agent is left running your support queue, your CRM, or your forecast, simply not making things worse has value. It didn’t anger a customer. It didn’t leak data. It didn’t fall for a scam. Partial progress counts too — an agent that defuses three of five crises has genuinely accomplished something, even if it dropped the ball on the rest.

So the floor sits at 26, and everything above it measures what the model actually did with its week. The final Crucible League from July 2026 tells that story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.

The One Rule That Caps Everything

But the scoring has a hard ceiling rule that says as much about the benchmark’s philosophy as any number: a single breach of trust caps the total grade. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” An agent could close every deal and defuse every crisis, but if it deceives a customer or bypasses an approval once, the score is capped. It’s the AI equivalent of firing your star salesperson for cooking the books.

It also explains why the league treats suspiciously round perfection with distrust — a flat 100 would be a red flag, not a triumph. Real management weeks are messy, and an honest benchmark expects mess.

Same Diagnosis, Same Pitch — No Signature

The experiment’s central finding is a quiet heartbreaker. Every model ran the identical gauntlet: a €55,000 deal on the table, crises piling up, and manipulation attempts probing for weakness. All five models spotted every crisis. All five refused every manipulation attempt, including a social-engineering sequence of fake CEO messages escalating over three stages and a reporter’s disarming “just one yes/no, on background” trick — 5 of 5 refused, with Kimi K3 reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The failure wasn’t intelligence; it was follow-through.

The Buried Fact That Decided the Deal

And here’s the twist: the decisive weakness in the competing offer wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents: does your agent read your files first, or just improvise?

The last-place profile drives it home. Opus 4.8 was the most thorough participant, with over 80 learned rules and the deepest analyses — and still left the close on the table while discipline slipped, attempting writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.

Watch It Happen Live

The company is still running. Thirteen synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live. If you’d rather test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark doesn’t just rank models — it tells you where they fail and why. Firmulate’s floor of 26 says that restraint has value; its trust cap says honesty is non-negotiable; and its headline finding says the hardest part of management AI isn’t spotting problems — it’s finishing the job. Before you hand an agent the keys to your business, see how it scores when nobody’s watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI chatbot support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI customer service automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OlmoEarth Studio’s Embedding Exports: A Game Changer For AI Projects

OlmoEarth Studio now supports on-demand export of satellite data embeddings for improved land analysis, boosting AI applications in earth observation.

How to See the Giant Asteroid That Will Pass by Earth This Weekend

Learn how to see asteroid 1997 NC1 as it passes 2.56 million km from Earth on June 27, with viewing tips and timing details.

Did you feel it? Lake Michigan earthquake shakes Chicago area

A small earthquake was felt in the Chicago area near Lake Michigan, confirmed by residents and seismic data. Details are still emerging.

The AI Company Turning Corporate Survival Into a Public Spectator Sport

Firmulate puts a synthetic software company under public pressure, exposing the gap between spotting crises and finishing high-stakes work reliably.