📊 Full opportunity report: Revealing AI’s Working Style Through A Strategic Management Assessment on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A recent experiment evaluated five AI management models during a simulated business crisis, revealing significant differences in their decision-making, trust, and action completion. The results highlight how AI can be assessed for real-world management tasks, as detailed in this analysis.
Five AI management models were tested in a live simulation managing a small software company during its worst week. The experiment revealed notable differences in how each model identified crises, maintained trust, and completed critical actions, providing new insights into AI’s management capabilities and limitations.
The experiment, conducted by Firmulate, involved five AI models running a simulated company with 13 synthetic employees, facing identical crises and pressures. For more details on the methodology, see the original analysis. Each model was tasked with making decisions that impacted the company’s cash flow, customer trust, and operational integrity. The models’ performance was scored based on their ability to diagnose problems, escalate risks appropriately, and close deals, with the top performer, gpt-5.6-sol, scoring 95 points out of 100.
Despite all models recognizing crises and refusing manipulative requests, only two successfully signed a critical €55,000 deal, which was the decisive step for the company’s revenue. The experiment also highlighted that thorough analysis alone did not guarantee success; effective action completion was crucial. This approach is similar to the management test that exposes an AI’s real working style. For example, Opus 4.8 produced detailed analyses but failed to close deals or escalate issues properly, illustrating that understanding must be paired with operational discipline.
Additionally, the models showed consistent strength in security instincts, refusing fake CEO requests designed to test manipulation, demonstrating their ability to recognize risks beyond obvious cues. The results underscore that AI’s management style varies significantly, and understanding these differences is essential for enterprise deployment.
Why AI Management Style Evaluation Matters
This experiment provides concrete evidence that AI models differ not only in their analytical depth but also in their ability to execute decisions, follow through, and maintain trust—critical factors for operational use. For businesses considering AI automation, these findings emphasize the importance of testing models in realistic scenarios before deployment. The ability to distinguish models that can identify crises, escalate appropriately, and complete deals directly impacts the effectiveness and safety of AI in management roles.
Furthermore, the results challenge assumptions that more analysis automatically leads to better management. Effective decision-making involves both understanding and action, a distinction made clear by the varied performances in this experiment. As AI becomes more integrated into business processes, such assessments will be vital for selecting models aligned with operational needs and risk management.

Principles of Agentic AI Governance: A Playbook for Managing AI Risk, Fairness, and Compliance (Agentic Governance and Architecture)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Context of AI Management Model Testing
Traditionally, AI tools have been evaluated based on their analytical and language capabilities. However, recent developments, including this live experiment conducted by Firmulate, focus on testing AI decision-making in realistic management scenarios. The experiment involved five frontier models running a simulated business through a crisis, with decisions recorded and scored in real time. This approach aims to bridge the gap between theoretical performance and practical operational readiness.
The experiment builds on prior efforts to understand AI’s role in management, emphasizing that trust, follow-through, and operational discipline are as important as analytical prowess. The models used—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—represent different AI frontiers, with results providing a benchmark for future evaluation and development.
“Testing AI models in realistic management scenarios reveals critical differences in their ability to execute and trustworthiness, which are often overlooked in standard benchmarks.”
— Firmulate
AI decision-making management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Performance
It is not yet clear how these results will translate to real-world enterprise environments, where variables and stakes are more complex. The experiment was conducted in a controlled simulation with synthetic employees and predefined crises, which may not capture all operational nuances. Additionally, the long-term reliability and trustworthiness of these models in ongoing management tasks remain to be tested in live settings.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Model Evaluation
Further testing is planned to evaluate these models in more complex, real-world scenarios. Enterprises are encouraged to run similar simulations tailored to their operations, using the same decision-recording framework to assess AI suitability. Development teams may also refine models to improve their follow-through and operational discipline, addressing the gaps identified in this experiment. The goal is to establish standardized benchmarks for AI management capabilities that go beyond analytical skills.

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment tell us about AI’s management abilities?
The experiment shows that AI models vary significantly in their ability to diagnose crises, maintain trust, and complete critical actions, which are essential for effective management.
Why is completing actions more important than analysis in management AI?
Because effective management requires not just understanding problems but also executing solutions, closing deals, and escalating issues when necessary.
Can these results predict how AI will perform in real businesses?
Not entirely. While the simulation provides valuable insights, real-world environments are more complex, and further testing is needed to confirm applicability.
What should companies do before deploying AI management models?
They should run tailored simulations and assessments to evaluate how models handle realistic crises, follow through on actions, and maintain trust.
Source: ThorstenMeyerAI.com