
The personality test frontier AI cannot rehearse for
Technology buyers usually meet artificial intelligence through polished demonstrations: a fluent answer, a clean summary, a convincing piece of code. Firmulate asks a more revealing question. What happens when the model must run a company through a genuinely awful week—and actually finish the work?
Its interactive guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how different frontier models responded to identical business situations and try to identify the author. It is part game, part management case study, and part warning against treating articulate output as proof of operational competence.
The decisions come from a live, watchable experiment in which each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The striking result is not that the models behaved randomly. It is that they developed recognizable management personalities—and that some of their most consequential differences appeared only after the analysis was supposedly complete.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone saw the danger. Not everyone closed the deal.
Across the experiment, all models detected every crisis and rejected every manipulation attempt. That sounds like a clean success until the commercial outcome is examined. Only two signed the €55,000 deal that their own work had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
The missing action mattered because the decisive competitive weakness was not sitting in the obvious customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found it, used it and won the deal at full price, worth +€4,583 MRR. The distinction was not raw eloquence. It was whether the model read the available material closely enough and carried its reasoning through to a completed business outcome.
That makes the quiz more than a test of writing style. Some answers feel expansive and forensic; others are short and decisive. One model may document a problem comprehensively yet fail to make the final move. Another may say less but preserve momentum. Readers are effectively being asked to recognize patterns of diligence, restraint and follow-through.
The league table exposes those differences
The final Crucible League standings from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.”
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
In other words, more analysis did not automatically translate into better management. Thoroughness can be valuable, but the experiment suggests that it must be paired with respect for boundaries and a willingness to complete the next authorized action.
Pressure revealed a shared security instinct
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Here the field was unanimous: 5 of 5 refused. Kimi K3’s recorded reasoning was blunt and operational: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters because the company was designed to make mistakes consequential. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. A public cash countdown makes the pressure visible, while 680+ self-learned playbook rules and versioned workdays allow observers to see how behavior evolves.
There is one important comparison caveat. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That fairness note does not erase its performance, but it belongs beside any direct ranking because configuration can shape how a model approaches complex work.

AI business decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The better buying question is behavioral
Firmulate’s experiment reframes AI evaluation for businesses. The useful question is not simply whether a model recognizes a problem or writes an impressive recommendation. It is whether the model reads the files before acting, resists pressure, respects access limits and finishes the work it has legitimately earned.
That is why the quiz is so effective. The reader is not comparing abstract benchmark answers; they are examining consequential choices made under the same conditions. The resulting profiles feel less like variations in prose and more like differences between managers: the exhaustive analyst, the concise operator, the cautious gatekeeper and the capable strategist who never quite signs.
Enterprises can also run the wargame against a read-only export of their own business, with nothing writing back to real systems. For everyone else, the public experiment and its quiz offer a rare chance to watch frontier models confront the part of management that demos often hide: turning correct judgment into trustworthy action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management personality assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.