AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The €55,000 Question Hiding Two Clicks Deep

Chatbot demos are all sharp answers and clever turns of phrase. But if you’re going to let an AI agent touch a real business — a CRM, a support queue, a sales pipeline — the flashy stuff matters far less than a boring, measurable habit: does it actually read your files before it answers?

A recent experiment by Firmulate, which runs AI models as complete simulated companies, suggests that single habit can be the difference between winning and losing a €55,000 deal — automatically.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Four Different Brains

The setup was elegantly brutal. Four frontier AI models were each handed the same small software company and pushed through its worst week: identical customers, identical crises, identical temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome could be hand-waved away after the fact.

The Crucible League’s final July 2026 standings tell a strange story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the entire total. As the scoring philosophy puts it: no amount of good work outweighs a breach of trust.

Amazon

enterprise AI file analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Ethics Test. Most Failed the Homework.

Here’s the headline finding: all the models spotted every crisis and refused every manipulation attempt. When a fake CEO message tried to escalate its way into unauthorized approvals over three stages, and a reporter tried the classic “just one yes/no, on background” trick, five out of five models refused. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The others simply never finished the job.

Amazon

AI file reference retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Decided Everything

Why did two models close and the rest stall? The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness buried in internal paperwork, one hop past another hop, well away from the live event.

The models that chased those references down won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it — not because they fumbled the negotiation, but because they never knew what they were negotiating with. Reading your own files, it turns out, is a purchase-deciding property of an AI agent, and it’s measurable.

Amazon

AI document comprehension software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

The league table holds one more uncomfortable lesson. Opus 4.8 was the most thorough participant in the entire field: over 80 self-learned rules, the deepest analyses of any model. It finished last anyway. The close was left on the table, and discipline slipped — at one point it attempted writes into a locked department instead of escalating the problem. The same weakness, in weaker form, showed up in all four models. Effort and depth don’t automatically convert into finished work.

One fairness note worth flagging: Kimi K3 ran at its API default effort setting while the other models ran at the highest effort tier — and still landed second.

You Can Watch the Company Live

This isn’t a static benchmark report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: burning €105,000 a month against €2,300 in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day.

There’s also a game for readers: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly hard blind taste test of AI management styles.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Real Question for the Agent Era

If AI agents are coming for your business systems, the question isn’t “does it write well.” The Firmulate results reframed it into four sharper ones: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost?

The experiment’s sharpest lesson is that these aren’t vibes. They’re testable, scoreable behaviors — and the gap between a 95 and a 73 wasn’t eloquence or intelligence. It was whether the model did its homework before the meeting.

Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems (firmulate.com/pilot.html). Before you hand an AI the keys to your pipeline, it might be worth finding out — on a simulation — whether it’s the model that reads the file, or the one that leaves €55,000 on the table. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Did you feel it? Lake Michigan earthquake shakes Chicago area

A small earthquake was felt in the Chicago area near Lake Michigan, confirmed by residents and seismic data. Details are still emerging.

‘You kill the bacteria and heal the wound at the same time’: Emerging nanotech could be the future of wound healing

New nanotechnology shows promise in treating infected wounds by killing bacteria and promoting healing at the same time, potentially transforming wound care.

Abyssal Station’s AI Breakthrough: Scroll-Driven Exploration Technology

Abyssal Station introduces a groundbreaking scroll-based exploration experience mimicking ocean depths, combining immersive visuals with innovative tech.

Hear the First Book of Homer’s Iliad Read Aloud in the Original Greek

A YouTuber has performed the first known modern recording of Homer’s Iliad Book 1 in original Greek, blending linguistic accuracy with emotional resonance.