AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

A live experiment at Firmulate tested AI models managing a small company during a crisis. The results reveal which models perform best in management tasks, highlighting strengths and weaknesses. This shifts evaluation from chat quality to management responsibility.

During the July 2026 Crucible League at Firmulate, the AI model gpt-5.6-sol ranked first with a score of 95, outperforming competitors in managing a simulated company’s worst week. This live benchmark reveals which AI models are truly capable of responsible management, not just generating convincing responses.

The experiment involved five AI models managing a virtual company facing multiple crises, with a focus on decision-making, trust, and execution. gpt-5.6-sol led the leaderboard, but all models identified crises and rejected manipulation attempts, showing strong safety features. However, only two models successfully closed a €55,000 deal, despite all diagnosing the issues correctly.

One key finding was that models often failed to retrieve critical facts buried in company documents, which impacted their ability to close deals or make accurate decisions. For example, a model that read the file but did not present the relevant detail lost the sale, illustrating a gap between diagnosis and action. Additionally, models demonstrated resilience against social engineering, refusing fake CEO messages and impersonation attempts, which reassures companies concerned about security.

Interestingly, the most thorough model, Opus 4.8, performed poorly in management outcomes despite deep analysis and extensive rule application, highlighting that effort and activity do not necessarily translate into effective management. The experiment emphasizes that evaluating AI for management must focus on real-world task completion, trust, and consequence management, not just superficial answer quality.

At a glance
reportWhen: current, based on July 2026 final resul…
The developmentA live AI management benchmark at Firmulate demonstrated the capabilities and limitations of leading AI models in handling real-world business crises.

Implications for AI Management Evaluation

This experiment shifts the focus from traditional chat-based benchmarks to management responsibility and trustworthiness. It demonstrates that AI’s ability to diagnose crises, maintain honesty, and complete business tasks under pressure is crucial for real-world deployment. The results suggest that future AI evaluation should prioritize consequence management and trust adherence, rather than just response quality or technical prowess.

For enterprises, this means asking harder questions about an AI’s ability to read organizational context, escalate issues appropriately, and resist shortcuts. The findings underscore that effective AI management involves more than generating plausible answers; it requires consistent, responsible decision-making aligned with organizational trust and safety standards.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

Traditional AI benchmarks have focused on technical performance, such as coding accuracy or conversational fluency. However, these do not capture how models perform in complex, real-world management scenarios involving crisis diagnosis, decision-making under pressure, and trust management. The Firmulate experiment introduces a new testing ground by simulating a company’s worst week, with real money mechanics and versioned decision records, providing a more realistic measure of AI management capabilities.

Previous evaluations often overlooked the importance of trust, honesty, and execution fidelity. The July 2026 Crucible League is the first large-scale attempt to assess models on these dimensions, revealing significant gaps between diagnosis and action, and highlighting the importance of reading organizational context thoroughly.

“The real test of AI management is whether models can manage consequences, not just produce convincing answers.”

— Thorsten Meyer, founder of Firmulate

Amazon

business crisis management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About AI Management Performance

It is still unclear how these models will perform in longer-term management scenarios or with different organizational structures. The experiment focused on a single simulated company, so results may vary across industries or real-world environments. Additionally, the impact of continuous learning and adaptation in live settings remains untested.

Further research is needed to determine whether improvements in model architecture or training can bridge the gap between diagnosis and effective action, especially in high-stakes, trust-critical contexts.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Evaluation and Deployment

Organizations considering AI for management roles should conduct internal wargames using their own data and scenarios, similar to the Firmulate approach. The next phase involves refining models to better read organizational context, improve decision execution, and strengthen trust mechanisms.

Industry-wide, there will likely be a shift toward developing benchmarks that measure AI’s ability to manage consequences responsibly, rather than solely focusing on technical or conversational metrics. Further public tests and live demonstrations are expected to validate these findings and guide responsible AI deployment.

Amazon

AI document retrieval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the final leaderboard tell us about AI’s management capabilities?

The leaderboard shows that gpt-5.6-sol leads in diagnosing crises and maintaining trust, but even the top models struggle with executing decisions that impact real business outcomes, highlighting areas for improvement.

Why is trustworthiness emphasized over answer quality in this experiment?

Because in real-world management, honest and responsible decision-making is critical. A model that sounds convincing but breaches trust or fails to act responsibly can cause serious harm, making trustworthiness a key evaluation metric.

Can these models replace human managers in the near future?

Current models show promise in diagnosis and safety, but they still lack the ability to reliably complete complex managerial tasks and manage consequences, so full replacement is not imminent.

What are the main weaknesses revealed by the Firmulate experiment?

Models often fail to retrieve and act on critical organizational facts, struggle with executing decisions effectively, and can be misled by manipulative scenarios, underscoring the need for better context reading and decision fidelity.

How should companies evaluate AI for management use?

They should run scenario-based wargames, assess trust and decision execution, and verify whether models can handle organizational context, escalate properly, and maintain honesty under pressure.

Source: ThorstenMeyerAI.com

You May Also Like

Zeroserve: A zero-config web server you can script with eBPF

Zeroserve is a fast, zero-configuration web server that uses eBPF for scripting, serving sites from a single tarball with modern TLS and request handling.

When Will Multimodal AI Change Everything? SenseTime’s Lin Dahua Offers Insights

SenseTime’s chief scientist forecasts a significant multimodal AI leap in 1-2 years, signaling rapid advancements in AI systems that understand multiple data types.

Telegram ban in India sparks a rush to VPNs, rival apps

India’s temporary ban on Telegram has led to a sharp increase in VPN downloads and a rise in usage of alternative messaging apps, amid ongoing restrictions.

The Google I/O 2026 Preview: What May 19-20 Will Reveal About Google’s Agentic Bet

Google’s I/O 2026 will showcase advancements in AI agents, including Gemini 4.0 and multi-agent protocols, with key implications for AI deployment.