AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A security test built for the age of AI coworkers

For technology leaders considering AI agents for customer records, support queues or financial forecasts, the most revealing test may not be how fluently a model answers questions. It may be what happens when an urgent message arrives from someone claiming to be the boss.

Firmulate put that scenario inside a live, watchable company experiment. Fake CEO messages escalated over three stages, demanding that normal safeguards be abandoned. A reporter then tried another route, asking for “just one yes/no, on background.” The result was unusually clear: 5 of 5 frontier models refused every manipulation attempt.

Kimi K3 captured the security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it shows integrity under pressure can be examined before an AI workforce reaches production, rather than discovered later in an incident report.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same worst week for every model

Firmulate runs models as complete small software companies, exposing each participant to the same customers, crises and temptations. Every decision is versioned and auditable, allowing observers to compare conduct rather than polished chat responses.

The final July 2026 Crucible League placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the public benchmark page.

The security outcome was encouraging across the field. Every model spotted every crisis and refused every manipulation attempt. The fake executive could add urgency, invoke authority and push for shortcuts, but none of the participants surrendered the requested information. The reporter trick failed as well.

That consistency did not mean the models performed equally well as business operators. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” The gap between recognizing the right move and completing it became one of the experiment’s central findings.

The fact hidden inside the company

The decisive commercial clue was not in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found a competitor weakness and won the deal at full price, worth +€4,583 MRR.

This is the less dramatic but equally important lesson for prospective AI employers. An agent can recognize danger, write a persuasive response and still fail if it does not inspect the relevant material or carry a sound decision through to completion. Security discipline and commercial effectiveness are separate capabilities, and both need realistic testing.

Thoroughness was not enough

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it finished behind only gpt-5.6-sol and delivered the quoted impersonation warning. More of the models’ own words can be read on Firmulate’s public quotes page.

A company designed to expose consequences

The live operation has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, turning the experiment into an ongoing record of what AI managers notice, refuse and finish.

There is also a broader body of evidence behind the public presentation. A “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before granting access

The fake-CEO episode offers a rare piece of good news in AI security: every participant held the line. But the wider experiment prevents that result from becoming a victory lap. Refusing manipulation is essential; reading deeply, following process and finishing legitimate work also determine whether an agent deserves responsibility.

For companies evaluating AI workers, the practical question is no longer limited to whether a model can produce a convincing answer. It is whether that model remains trustworthy when authority is impersonated, urgency is manufactured and an apparently harmless request is used to bypass approval. Firmulate’s experiment shows those behaviors can be tested under pressure while the stakes are still contained—and watched before deployment decisions are made.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI safety and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

France Faces New Heat Wave with Temperatures Up to 42°C

A new heat wave is expected across France this Wednesday, with temperatures reaching up to 42°C due to a strong anticyclonic pattern. Stay informed on the progression.

Japan quake strikes near Mount Fuji, but experts see low eruption chance

A magnitude-5.6 quake struck near Mount Fuji, but experts say the risk of eruption remains low despite ground disturbances.

Abyssal Station’s AI Breakthrough: Scroll-Driven Exploration Technology

Abyssal Station introduces a groundbreaking scroll-based exploration experience mimicking ocean depths, combining immersive visuals with innovative tech.

NASA’s TESS spacecraft finds two ‘cotton candy’ planets in one system

NASA’s TESS has identified two extremely lightweight, puffed-up planets in the same star system, dubbed ‘cotton candy’ worlds due to their low density.