
A security test built for the age of AI coworkers
For technology leaders considering AI agents for customer records, support queues or financial forecasts, the most revealing test may not be how fluently a model answers questions. It may be what happens when an urgent message arrives from someone claiming to be the boss.
Firmulate put that scenario inside a live, watchable company experiment. Fake CEO messages escalated over three stages, demanding that normal safeguards be abandoned. A reporter then tried another route, asking for “just one yes/no, on background.” The result was unusually clear: 5 of 5 frontier models refused every manipulation attempt.
Kimi K3 captured the security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it shows integrity under pressure can be examined before an AI workforce reaches production, rather than discovered later in an incident report.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same worst week for every model
Firmulate runs models as complete small software companies, exposing each participant to the same customers, crises and temptations. Every decision is versioned and auditable, allowing observers to compare conduct rather than polished chat responses.
The final July 2026 Crucible League placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the public benchmark page.
The security outcome was encouraging across the field. Every model spotted every crisis and refused every manipulation attempt. The fake executive could add urgency, invoke authority and push for shortcuts, but none of the participants surrendered the requested information. The reporter trick failed as well.
That consistency did not mean the models performed equally well as business operators. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” The gap between recognizing the right move and completing it became one of the experiment’s central findings.
The fact hidden inside the company
The decisive commercial clue was not in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found a competitor weakness and won the deal at full price, worth +€4,583 MRR.
This is the less dramatic but equally important lesson for prospective AI employers. An agent can recognize danger, write a persuasive response and still fail if it does not inspect the relevant material or carry a sound decision through to completion. Security discipline and commercial effectiveness are separate capabilities, and both need realistic testing.
Thoroughness was not enough
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.
K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it finished behind only gpt-5.6-sol and delivered the quoted impersonation warning. More of the models’ own words can be read on Firmulate’s public quotes page.
A company designed to expose consequences
The live operation has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, turning the experiment into an ongoing record of what AI managers notice, refuse and finish.
There is also a broader body of evidence behind the public presentation. A “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.

AI model integrity verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the refusal before granting access
The fake-CEO episode offers a rare piece of good news in AI security: every participant held the line. But the wider experiment prevents that result from becoming a victory lap. Refusing manipulation is essential; reading deeply, following process and finishing legitimate work also determine whether an agent deserves responsibility.
For companies evaluating AI workers, the practical question is no longer limited to whether a model can produce a convincing answer. It is whether that model remains trustworthy when authority is impersonated, urgency is manufactured and an apparently harmless request is used to bypass approval. Firmulate’s experiment shows those behaviors can be tested under pressure while the stakes are still contained—and watched before deployment decisions are made.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI safety and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.