Revealing AI’s Working Style Through A Strategic Management Assessment
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Revealing AI’s Working Style Through A Strategic Management Assessment on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A recent experiment evaluated five AI management models during a simulated business crisis, revealing significant differences in their decision-making, trust, and action completion. The results highlight how AI can be assessed for real-world management tasks, as detailed in this analysis.

Five AI management models were tested in a live simulation managing a small software company during its worst week. The experiment revealed notable differences in how each model identified crises, maintained trust, and completed critical actions, providing new insights into AI’s management capabilities and limitations.

The experiment, conducted by Firmulate, involved five AI models running a simulated company with 13 synthetic employees, facing identical crises and pressures. For more details on the methodology, see the original analysis. Each model was tasked with making decisions that impacted the company’s cash flow, customer trust, and operational integrity. The models’ performance was scored based on their ability to diagnose problems, escalate risks appropriately, and close deals, with the top performer, gpt-5.6-sol, scoring 95 points out of 100.

Despite all models recognizing crises and refusing manipulative requests, only two successfully signed a critical €55,000 deal, which was the decisive step for the company’s revenue. The experiment also highlighted that thorough analysis alone did not guarantee success; effective action completion was crucial. This approach is similar to the management test that exposes an AI’s real working style. For example, Opus 4.8 produced detailed analyses but failed to close deals or escalate issues properly, illustrating that understanding must be paired with operational discipline.

Additionally, the models showed consistent strength in security instincts, refusing fake CEO requests designed to test manipulation, demonstrating their ability to recognize risks beyond obvious cues. The results underscore that AI’s management style varies significantly, and understanding these differences is essential for enterprise deployment.

At a glance
reportWhen: ongoing; results announced July 2026
The developmentA live experiment tested AI models’ ability to manage a small software company’s worst week, exposing their decision-making and operational behaviors.

Why AI Management Style Evaluation Matters

This experiment provides concrete evidence that AI models differ not only in their analytical depth but also in their ability to execute decisions, follow through, and maintain trust—critical factors for operational use. For businesses considering AI automation, these findings emphasize the importance of testing models in realistic scenarios before deployment. The ability to distinguish models that can identify crises, escalate appropriately, and complete deals directly impacts the effectiveness and safety of AI in management roles.

Furthermore, the results challenge assumptions that more analysis automatically leads to better management. Effective decision-making involves both understanding and action, a distinction made clear by the varied performances in this experiment. As AI becomes more integrated into business processes, such assessments will be vital for selecting models aligned with operational needs and risk management.

Principles of Agentic AI Governance: A Playbook for Managing AI Risk, Fairness, and Compliance (Agentic Governance and Architecture)

Principles of Agentic AI Governance: A Playbook for Managing AI Risk, Fairness, and Compliance (Agentic Governance and Architecture)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Context of AI Management Model Testing

Traditionally, AI tools have been evaluated based on their analytical and language capabilities. However, recent developments, including this live experiment conducted by Firmulate, focus on testing AI decision-making in realistic management scenarios. The experiment involved five frontier models running a simulated business through a crisis, with decisions recorded and scored in real time. This approach aims to bridge the gap between theoretical performance and practical operational readiness.

The experiment builds on prior efforts to understand AI’s role in management, emphasizing that trust, follow-through, and operational discipline are as important as analytical prowess. The models used—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—represent different AI frontiers, with results providing a benchmark for future evaluation and development.

“Testing AI models in realistic management scenarios reveals critical differences in their ability to execute and trustworthiness, which are often overlooked in standard benchmarks.”

— Firmulate

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Performance

It is not yet clear how these results will translate to real-world enterprise environments, where variables and stakes are more complex. The experiment was conducted in a controlled simulation with synthetic employees and predefined crises, which may not capture all operational nuances. Additionally, the long-term reliability and trustworthiness of these models in ongoing management tasks remain to be tested in live settings.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Model Evaluation

Further testing is planned to evaluate these models in more complex, real-world scenarios. Enterprises are encouraged to run similar simulations tailored to their operations, using the same decision-recording framework to assess AI suitability. Development teams may also refine models to improve their follow-through and operational discipline, addressing the gaps identified in this experiment. The goal is to establish standardized benchmarks for AI management capabilities that go beyond analytical skills.

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment tell us about AI’s management abilities?

The experiment shows that AI models vary significantly in their ability to diagnose crises, maintain trust, and complete critical actions, which are essential for effective management.

Why is completing actions more important than analysis in management AI?

Because effective management requires not just understanding problems but also executing solutions, closing deals, and escalating issues when necessary.

Can these results predict how AI will perform in real businesses?

Not entirely. While the simulation provides valuable insights, real-world environments are more complex, and further testing is needed to confirm applicability.

What should companies do before deploying AI management models?

They should run tailored simulations and assessments to evaluate how models handle realistic crises, follow through on actions, and maintain trust.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The Enforcement Countdown: 89 Days Until the EU AI Act’s GPAI Penalty Phase Begins

In 89 days, the EU will activate enforcement powers under the AI Act for GPAI providers, enabling fines up to €35M or 7% of turnover. Companies must prepare now.

VigilSAR Benchmark: There Is No Best Model

VigilSAR’s new benchmark shows no model excels across all defense-relevant axes, emphasizing context-specific model selection over a universal leader.

Preparing Your Company For 2026: OpenAI’s Data Strategy For AI Success

OpenAI announces a comprehensive data governance approach for enterprise AI products, emphasizing privacy, control, and security ahead of 2026.

The Menu: What Ten Answers Reveal

A detailed analysis of how ten jurisdictions respond to automation, AI, and income risks, revealing patterns, strengths, and political instincts.