📊 Full opportunity report: AI Management Test: A Window Into Its Real Working Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
An ongoing AI management experiment pits five frontier models against a simulated company’s worst week. The models’ ability to diagnose, act, and complete crucial tasks reveals notable differences in discipline and follow-through, offering insights into AI’s practical management capabilities.
A live experiment on firmulate.com is testing five AI management models by having them run a simulated company through its worst week. The models are evaluated on their ability to diagnose crises, maintain trust, and complete critical actions, providing a rare, real-time look into AI’s practical management capabilities. This development matters because it exposes the actual operational strengths and weaknesses of AI in business decision-making, as detailed in the original analysis.
The experiment, called the Crucible League, involves five frontier AI models managing a small software company with a monthly burn rate of €105,000 against €2,300 in recurring revenue. The models are tasked with handling identical crises, customer issues, and business pressures, with their decisions recorded and auditable. The models’ performance is scored based on their ability to identify problems, act decisively, and follow through to close deals or escalate issues appropriately.
Results from July 2026 show gpt-5.6-sol leading with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline model scored only 26, highlighting the gap between simple analysis and effective action. Despite all models recognizing crises and refusing manipulation attempts, only two successfully closed deals, underscoring that diagnosis alone is insufficient for management success.
One notable finding is that thorough analysis does not guarantee operational discipline. Opus 4.8, despite producing the deepest insights, failed to close a key deal due to operational lapses, illustrating that effective management requires both understanding and execution. The experiment also emphasizes that AI models can recognize risks, such as social engineering attempts, but struggle with executing the final steps needed to achieve business outcomes.
Implications for AI in Business Management
This experiment demonstrates that AI’s value in management depends not just on analysis but on execution. The ability to identify crises, maintain trust, and follow through on decisions is crucial for AI to be genuinely useful in operational roles. The results suggest that enterprises should evaluate AI models based on their real-world decision-making and action-taking capabilities, not just their analytical depth. This approach helps prevent overestimating AI’s readiness for autonomous management and highlights the importance of testing AI in realistic scenarios before deployment.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing and Its Evolution
Traditional AI demonstrations often focus on theoretical or simulated tasks without revealing how models perform under real operational pressures. Firmulate’s live experiment is a departure, providing a transparent, auditable environment where AI models are tested against actual business crises. Prior to this, AI assessments mainly measured language proficiency or problem-solving in controlled settings. This experiment extends evaluation into the realm of practical management, emphasizing decision follow-through and operational discipline, which are critical for AI to be trusted as a management tool.
“Same diagnosis, same pitch — no signature.”
— Firmulate.com summary

Social Simulation for a Crisis: Results and Lessons from Simulating the COVID-19 Crisis (Computational Social Sciences)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Effectiveness
It remains unclear how these models will perform over longer periods or in different types of business environments. The experiment focuses on a single simulated scenario, and real-world complexities could produce different results. Additionally, the impact of different operational parameters, such as effort settings or integration with actual business systems, needs further exploration. The extent to which these findings generalize to live enterprise management is still under investigation.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Steps for Testing and Deploying AI Management Models
The next phase involves expanding the experiment to include more diverse scenarios and longer management periods. Enterprises are encouraged to replicate similar tests using their own business data in controlled environments, assessing how AI models handle real operational pressures. Further research will explore improving models’ operational discipline and decision follow-through, aiming to bridge the gap between analysis and action. Results from subsequent league rounds and real-world pilot projects are expected in late 2026.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the purpose of the Firmulate AI management experiment?
The experiment aims to evaluate how well AI models can manage real business crises, focusing on their decision-making, trust preservation, and ability to complete critical actions in a simulated company environment.
Which AI models performed best in the experiment?
According to the July 2026 results, gpt-5.6-sol scored highest, followed closely by Kimi K3. Both demonstrated stronger operational discipline and decision follow-through than other models.
Why is operational discipline more important than analysis in AI management?
The experiment shows that even the most thorough analysis does not guarantee successful management if the AI fails to act decisively and follow through on critical steps, such as closing deals or escalating issues properly.
Can these AI models replace human managers?
Not yet. The experiment highlights that AI still struggles with translating diagnosis into action consistently. Further development and testing are needed before AI can be considered a reliable replacement for human management.
What are the next steps for businesses interested in AI management tools?
Businesses should consider running their own controlled tests using their operational scenarios to evaluate AI models’ decision-making and execution capabilities before deploying them in live environments.
Source: ThorstenMeyerAI.com