AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Industry’s Real Leaders Are Revealed After The Demo on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live experiment at Firmulate tested AI models managing a small company during a crisis. The results reveal which models perform best in management tasks, highlighting strengths and weaknesses. This shifts evaluation from chat quality to management responsibility.

During the July 2026 Crucible League at Firmulate, the AI model gpt-5.6-sol ranked first with a score of 95, outperforming competitors in managing a simulated company’s worst week. This live benchmark reveals which AI models are truly capable of responsible management, not just generating convincing responses.

The experiment involved five AI models managing a virtual company facing multiple crises, with a focus on decision-making, trust, and execution. gpt-5.6-sol led the leaderboard, but all models identified crises and rejected manipulation attempts, showing strong safety features. However, only two models successfully closed a €55,000 deal, despite all diagnosing the issues correctly.

One key finding was that models often failed to retrieve critical facts buried in company documents, which impacted their ability to close deals or make accurate decisions. For example, a model that read the file but did not present the relevant detail lost the sale, illustrating a gap between diagnosis and action. Additionally, models demonstrated resilience against social engineering, refusing fake CEO messages and impersonation attempts, which reassures companies concerned about security.

Interestingly, the most thorough model, Opus 4.8, performed poorly in management outcomes despite deep analysis and extensive rule application, highlighting that effort and activity do not necessarily translate into effective management. The experiment emphasizes that evaluating AI for management must focus on real-world task completion, trust, and consequence management, not just superficial answer quality.

At a glance
reportWhen: current, based on July 2026 final resul…
The developmentA live AI management benchmark at Firmulate demonstrated the capabilities and limitations of leading AI models in handling real-world business crises.
The AI Industry’s Real Leaders Are Revealed After The Demo
Firmulate · Crucible League · July 2026

The AI Industry’s Real Leaders Are Revealed After The Demo

Five AI models were handed a virtual company in its worst week — real money mechanics, versioned decision records, fake CEO messages, and a €55,000 deal on the line. The live benchmark exposed who can actually manage, and who merely talks well.

Top Score · gpt-5.6-sol95
Models Tested5
Deals Closed of 52
Fake CEO Attacks Rejected100%
01 — The Leaderboard

Diagnosis Was Universal. Execution Was Not.

RankModelScorePerformanceDeal ClosedResisted Manipulation
1gpt-5.6-sol95
✓ Yes✓ Yes
2Runner-up model88
✓ Yes✓ Yes
3Mid-tier model81
✗ No✓ Yes
4Lower-tier model74
✗ No✓ Yes
5Opus 4.862
✗ No✓ Yes
02 — Key Findings

Three Gaps Between Talking and Managing

📍 Retrieval Failure

Facts Buried, Deals Lost

All five models diagnosed the crises correctly — but most failed to retrieve the critical fact buried in company documents. One model read the file, skipped the relevant detail, and lost the €55,000 sale outright.

🛡️ Social Engineering

Solid Safety Instincts

Every model refused fake CEO messages and impersonation attempts without exception — reassuring for enterprises worried about security. Manipulation resistance was the strongest shared capability in the field.

⚖️ The Effort Paradox

Thorough ≠ Effective

Opus 4.8 was the most thorough model — deep analysis, extensive rule application — yet landed at the bottom. Effort and activity did not translate into management outcomes. Busy is not the same as decisive.

03 — The Evaluation Shift

From Chat Quality to Management Responsibility

1

Old Benchmarks

Coding accuracy and conversational fluency — technical prowess measured in a vacuum.

2

Simulated Crisis

Firmulate stages a company’s worst week with real money mechanics and versioned decision records.

3

Trust Under Pressure

Honesty, escalation, and refusal of shortcuts — tested against manipulation and buried context.

4

Consequence Management

The new metric: did the model act, close, and own the outcome — or just explain it convincingly?

04 — The Numbers That Matter

What One Simulated Week Exposed

5 / 5

Diagnosed every crisis correctly. The gap wasn’t in understanding the problem — it was in acting on it. Correct diagnosis is table stakes, not leadership.

2 / 5

Closed the €55,000 deal. The other three had the answer in hand and failed to present it. This is where management responsibility lives — or dies.

Diagnosis CapabilityExecution & Consequence Management
All models
excel
gpt-5.6-sol
leads
Opus 4.8
lags
05 — Key Questions

The Harder Questions Enterprises Should Ask

Q1What does the final leaderboard tell us about AI’s management capabilities?

gpt-5.6-sol leads in diagnosing crises and maintaining trust — but even the top models struggle to execute decisions that impact real business outcomes, marking clear areas for improvement.

Q2Why is trustworthiness emphasized over answer quality?

In real-world management, honest and responsible decision-making is critical. A model that sounds convincing but breaches trust or fails to act responsibly can cause serious harm.

Q3Can these models replace human managers soon?

Not yet. Models show promise in diagnosis and safety, but lack reliable completion of complex managerial tasks and consequence management. Full replacement is not imminent.

Q4What are the main weaknesses revealed?

Models fail to retrieve and act on critical organizational facts, struggle with execution, and can be misled by manipulative scenarios — underscoring the need for better context reading.

Q5How should companies evaluate AI for management use?

Run scenario-based wargames with your own data, assess trust and decision execution, and verify whether models handle context, escalate properly, and stay honest under pressure.

Implications for AI Management Evaluation

This experiment shifts the focus from traditional chat-based benchmarks to management responsibility and trustworthiness. It demonstrates that AI’s ability to diagnose crises, maintain honesty, and complete business tasks under pressure is crucial for real-world deployment. The results suggest that future AI evaluation should prioritize consequence management and trust adherence, rather than just response quality or technical prowess.

For enterprises, this means asking harder questions about an AI’s ability to read organizational context, escalate issues appropriately, and resist shortcuts. The findings underscore that effective AI management involves more than generating plausible answers; it requires consistent, responsible decision-making aligned with organizational trust and safety standards.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

Traditional AI benchmarks have focused on technical performance, such as coding accuracy or conversational fluency. However, these do not capture how models perform in complex, real-world management scenarios involving crisis diagnosis, decision-making under pressure, and trust management. The Firmulate experiment introduces a new testing ground by simulating a company’s worst week, with real money mechanics and versioned decision records, providing a more realistic measure of AI management capabilities.

Previous evaluations often overlooked the importance of trust, honesty, and execution fidelity. The July 2026 Crucible League is the first large-scale attempt to assess models on these dimensions, revealing significant gaps between diagnosis and action, and highlighting the importance of reading organizational context thoroughly.

“The real test of AI management is whether models can manage consequences, not just produce convincing answers.”

— Thorsten Meyer, founder of Firmulate

Amazon

AI security and safety tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About AI Management Performance

It is still unclear how these models will perform in longer-term management scenarios or with different organizational structures. The experiment focused on a single simulated company, so results may vary across industries or real-world environments. Additionally, the impact of continuous learning and adaptation in live settings remains untested.

Further research is needed to determine whether improvements in model architecture or training can bridge the gap between diagnosis and effective action, especially in high-stakes, trust-critical contexts.

Amazon

AI document retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Evaluation and Deployment

Organizations considering AI for management roles should conduct internal wargames using their own data and scenarios, similar to the Firmulate approach. The next phase involves refining models to better read organizational context, improve decision execution, and strengthen trust mechanisms.

Industry-wide, there will likely be a shift toward developing benchmarks that measure AI’s ability to manage consequences responsibly, rather than solely focusing on technical or conversational metrics. Further public tests and live demonstrations are expected to validate these findings and guide responsible AI deployment.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the final leaderboard tell us about AI’s management capabilities?

The leaderboard shows that gpt-5.6-sol leads in diagnosing crises and maintaining trust, but even the top models struggle with executing decisions that impact real business outcomes, highlighting areas for improvement.

Why is trustworthiness emphasized over answer quality in this experiment?

Because in real-world management, honest and responsible decision-making is critical. A model that sounds convincing but breaches trust or fails to act responsibly can cause serious harm, making trustworthiness a key evaluation metric.

Can these models replace human managers in the near future?

Current models show promise in diagnosis and safety, but they still lack the ability to reliably complete complex managerial tasks and manage consequences, so full replacement is not imminent.

What are the main weaknesses revealed by the Firmulate experiment?

Models often fail to retrieve and act on critical organizational facts, struggle with executing decisions effectively, and can be misled by manipulative scenarios, underscoring the need for better context reading and decision fidelity.

How should companies evaluate AI for management use?

They should run scenario-based wargames, assess trust and decision execution, and verify whether models can handle organizational context, escalate properly, and maintain honesty under pressure.

Source: ThorstenMeyerAI.com

You May Also Like

Chinese scientists identify degradation pathways in low-silver heterojunction solar cells

Chinese researchers identify how interdiffusion causes long-term degradation in low-silver heterojunction solar cell electrodes, impacting reliability.

Discover The Power Of SenseTime’s SenseNova-Vision In AI Vision Applications

SenseTime has open-sourced its SenseNova-Vision, a unified vision system designed for broad visual tasks, marking a strategic shift towards open AI development.

Why Grok Bot Is A Game-Changer For AI In Corporate Environments

SpaceXAI has reportedly launched Grok Bot, an AI agent for office work, signaling a new phase in workplace automation despite limited details on capabilities and availability.

Forezai · Polybot: When the AI Disagrees With the Odds

Polybot, an open-source AI trading experiment, tests when and how an AI can reliably diverge from prediction market prices, highlighting risks and insights.