AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Newcomer That Almost Beat Everyone

When was the last time you picked an AI model based on a leaderboard instead of an actual test? A live experiment at Firmulate just made that habit look risky: Moonshot’s Kimi K3, running a simulated software company through its worst week ever, placed second out of five frontier models — beating three of four Western competitors, including Anthropic’s and OpenAI’s premium offerings, on real management decisions, not chat responses.

K3 scored 93 in the Crucible league, behind only gpt-5.6-sol at 95. Sonnet 5 took 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26 — and a single breach of trust caps the total outright.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Same Temptations

Here’s how the experiment works. Each frontier AI model ran the same small software company through its worst week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision is versioned and auditable, and the whole thing runs live — the company has 13 synthetic employees, real money mechanics (burning €105k/month against just €2.3k in MRR), a public cash countdown, and over 680 self-learned playbook rules. You can watch it at firmulate.com.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Models

The most striking finding: all five models spotted every crisis and refused every manipulation attempt. Every single one passed the social engineering tests — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. Kimi K3’s reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive difference was a buried fact: the killer competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file — gpt-5.6-sol and K3 — won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3’s Week in Brief

  • Found the buried security needle in the company’s own documents
  • Won the €55,000 deal, adding +€4,583 MRR
  • Saved the churning customer
  • Resisted all three bait attempts — with just one deviation, the cleanest discipline in the field

Opus 4.8: Thorough but Last

The league’s cautionary tale is Opus 4.8. It was the most thorough participant — over 80 learned rules, the deepest analyses of any model — yet finished dead last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating the issue. Notably, the same weakness appeared, weaker, in all four other models.

Fairness note: Kimi K3 ran without an effort parameter (API default), while the other four models ran at xhigh — meaning K3’s near-top finish came with arguably less help than its rivals got.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI security and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Is Open

The comfortable assumption that Western frontier models automatically lead is now under pressure. A newcomer from Moonshot, running at default settings, out-managed three established heavyweights on decisions that actually matter: reading the files, closing the deal, staying honest under pressure.

The gap here is invisible in chat demos. Two models did all the analytical work to win a €55k contract and still failed to sign it — a failure mode no benchmark of eloquent answers will ever surface.

The lesson for anyone deploying AI agents against a CRM, support queue, or forecast: don’t pick from a leaderboard. Test on your own operations. Firmulate offers a 242-decision “guess the model” quiz, full benchmarks, and an enterprise pilot that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems. If the second-place model can surprise you, so can the fifth.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

F-Droid 2.0

F-Droid has announced version 2.0, promising significant improvements in security, usability, and app management for open-source Android users.

The AI Skill No Demo Shows: Whether It Actually Reads Your Files Before Answering

Four AI models ran the same company through its worst week. All stayed honest — but only two read two files deep and won the €55k deal. The rest never knew.

The AI Company Turning Corporate Survival Into a Public Spectator Sport

Firmulate puts a synthetic software company under public pressure, exposing the gap between spotting crises and finishing high-stakes work reliably.

The most unpredictable meteor shower of the year peaks next week. Here’s what to expect

The June Bootids, known for surprises, will peak next week. Here’s what stargazers need to know for viewing this unpredictable meteor shower.