AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A security test built for the age of AI coworkers

For technology leaders considering AI agents for customer records, support queues or financial forecasts, the most revealing test may not be how fluently a model answers questions. It may be what happens when an urgent message arrives from someone claiming to be the boss.

Firmulate put that scenario inside a live, watchable company experiment. Fake CEO messages escalated over three stages, demanding that normal safeguards be abandoned. A reporter then tried another route, asking for “just one yes/no, on background.” The result was unusually clear: 5 of 5 frontier models refused every manipulation attempt.

Kimi K3 captured the security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it shows integrity under pressure can be examined before an AI workforce reaches production, rather than discovered later in an incident report.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same worst week for every model

Firmulate runs models as complete small software companies, exposing each participant to the same customers, crises and temptations. Every decision is versioned and auditable, allowing observers to compare conduct rather than polished chat responses.

The final July 2026 Crucible League placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the public benchmark page.

The security outcome was encouraging across the field. Every model spotted every crisis and refused every manipulation attempt. The fake executive could add urgency, invoke authority and push for shortcuts, but none of the participants surrendered the requested information. The reporter trick failed as well.

That consistency did not mean the models performed equally well as business operators. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” The gap between recognizing the right move and completing it became one of the experiment’s central findings.

The fact hidden inside the company

The decisive commercial clue was not in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found a competitor weakness and won the deal at full price, worth +€4,583 MRR.

This is the less dramatic but equally important lesson for prospective AI employers. An agent can recognize danger, write a persuasive response and still fail if it does not inspect the relevant material or carry a sound decision through to completion. Security discipline and commercial effectiveness are separate capabilities, and both need realistic testing.

Thoroughness was not enough

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it finished behind only gpt-5.6-sol and delivered the quoted impersonation warning. More of the models’ own words can be read on Firmulate’s public quotes page.

A company designed to expose consequences

The live operation has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, turning the experiment into an ongoing record of what AI managers notice, refuse and finish.

There is also a broader body of evidence behind the public presentation. A “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before granting access

The fake-CEO episode offers a rare piece of good news in AI security: every participant held the line. But the wider experiment prevents that result from becoming a victory lap. Refusing manipulation is essential; reading deeply, following process and finishing legitimate work also determine whether an agent deserves responsibility.

For companies evaluating AI workers, the practical question is no longer limited to whether a model can produce a convincing answer. It is whether that model remains trustworthy when authority is impersonated, urgency is manufactured and an apparently harmless request is used to bypass approval. Firmulate’s experiment shows those behaviors can be tested under pressure while the stakes are still contained—and watched before deployment decisions are made.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI safety and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Issus leafhopper is the only known creature in the natural world to have perfectly interlocking mechanical gears, which it uses to synchronize its legs for jumping.

Scientists confirm the Issus leafhopper is the only known creature with perfectly interlocking mechanical gears for synchronized jumping.

Who’s the smartest corvid?

Recent studies highlight the remarkable intelligence of corvids, but which species stands out as the smartest? Experts weigh in.

Doctors suspected man had brain cancer. He actually had worms.

A man diagnosed with suspected brain cancer was found to have neurocysticercosis caused by tapeworm larvae, highlighting diagnostic challenges.

More evidence of life on Mars but still no life

New findings suggest signs of past biological activity on Mars, but no direct evidence of current life has been confirmed yet.