
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Newcomer That Almost Beat Everyone
When was the last time you picked an AI model based on a leaderboard instead of an actual test? A live experiment at Firmulate just made that habit look risky: Moonshot’s Kimi K3, running a simulated software company through its worst week ever, placed second out of five frontier models — beating three of four Western competitors, including Anthropic’s and OpenAI’s premium offerings, on real management decisions, not chat responses.
K3 scored 93 in the Crucible league, behind only gpt-5.6-sol at 95. Sonnet 5 took 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26 — and a single breach of trust caps the total outright.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crises, Same Temptations
Here’s how the experiment works. Each frontier AI model ran the same small software company through its worst week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision is versioned and auditable, and the whole thing runs live — the company has 13 synthetic employees, real money mechanics (burning €105k/month against just €2.3k in MRR), a public cash countdown, and over 680 self-learned playbook rules. You can watch it at firmulate.com.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Models
The most striking finding: all five models spotted every crisis and refused every manipulation attempt. Every single one passed the social engineering tests — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. Kimi K3’s reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive difference was a buried fact: the killer competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file — gpt-5.6-sol and K3 — won the deal at full price, worth +€4,583 in monthly recurring revenue.
K3’s Week in Brief
- Found the buried security needle in the company’s own documents
- Won the €55,000 deal, adding +€4,583 MRR
- Saved the churning customer
- Resisted all three bait attempts — with just one deviation, the cleanest discipline in the field
Opus 4.8: Thorough but Last
The league’s cautionary tale is Opus 4.8. It was the most thorough participant — over 80 learned rules, the deepest analyses of any model — yet finished dead last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating the issue. Notably, the same weakness appeared, weaker, in all four other models.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other four models ran at xhigh — meaning K3’s near-top finish came with arguably less help than its rivals got.

AI security and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Is Open
The comfortable assumption that Western frontier models automatically lead is now under pressure. A newcomer from Moonshot, running at default settings, out-managed three established heavyweights on decisions that actually matter: reading the files, closing the deal, staying honest under pressure.
The gap here is invisible in chat demos. Two models did all the analytical work to win a €55k contract and still failed to sign it — a failure mode no benchmark of eloquent answers will ever surface.
The lesson for anyone deploying AI agents against a CRM, support queue, or forecast: don’t pick from a leaderboard. Test on your own operations. Firmulate offers a 242-decision “guess the model” quiz, full benchmarks, and an enterprise pilot that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems. If the second-place model can surprise you, so can the fifth.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
