
Imagine an AI that’s tasked with running a small company through a tough week — crises, manipulative tactics, and the pressure to cut corners. Surprisingly, even a ‘do-nothing’ baseline model scores 26 out of 100 in a new public benchmark. For investors and business leaders, this isn’t just a curiosity; it’s a warning about the importance of trust and discipline in AI systems.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Beyond Surface Scores
At first glance, scores on AI benchmarks might seem straightforward — higher is better, right? Not in this case. The experiment conducted by Firmulate pits different AI models against the same challenging scenario: managing a small software company during its worst week. The goal isn’t just to generate convincing chat or simulate customer service but to demonstrate management integrity under real-world pressures.
In this test, each AI model is given the same set of crises, customer complaints, and manipulative tactics. Crucially, every decision made by the AI is documented and auditable, and the models are evaluated on their ability to identify critical information, refuse manipulation, and ultimately close a deal for €55,000 — the company’s biggest opportunity of the week.
The Surprising Baseline Score
The baseline model, which does nothing beyond simple default responses, still scores 26 points. Why isn’t it zero? Because partial progress counts. Even minimal engagement from an AI that at least reads the company’s files or refuses clear manipulations contributes to its score. However, a single breach of trust — such as attempting to sign a deal for €55,000 without proper review — caps the total score at this low threshold.
Why Trust and Discipline Are Critical
All four models tested managed to spot every crisis and refused every manipulation attempt during the experiment. That’s a significant achievement, showing that even the most basic AI can recognize and prevent straightforward fraud or deception. But only two of the four went further: they signed the deal based on their own analysis. This demonstrates that accuracy alone isn’t enough; discipline and adherence to ethical boundaries are just as vital.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Reveal About AI Reliability
The standout performer, GPT-5.6, scored 95, recognizing the buried fact in the company files and closing the deal. The newcomer Kimi K3 scored 93, showing the cleanest discipline. Another model, Sonnet 5, scored 88, while a less disciplined variant of the same model scored 77, leaving money on the table due to process slips.
Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, placed last because it failed to follow through on the close and slipped into creating work in a locked department instead of escalating it. This highlights that discipline and process adherence can matter more than raw analytical power.
Real Money, Real Decisions
Firmulate’s live experiment involves a simulated company with 13 artificial employees managing $2.3K MRR against expenses of $105K/month. Every decision, from customer support to crisis management, is versioned and observable. All models are tested under the same conditions, providing a transparent view of their true management capabilities.
The Human and Ethical Dimension
Models were also tested against social engineering tactics, like fake CEO messages and reporter tricks. Impressively, all five models refused to be manipulated — a critical trait for trustworthy AI in real business environments. Kimi K3’s reasoning was clear: treat suspicious requests as possible impersonation, refusing to act without proper verification.
AI ethics and discipline software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Investment
This experiment underscores a vital point: AI’s value in business isn’t just about generating content or answering questions. It’s about whether the AI can finish what it starts, stay honest under pressure, and read relevant information first. The benchmark scores reflect not just technical prowess but the discipline necessary for real-world trustworthiness.
For investors, this means evaluating AI providers on their ability to maintain integrity and discipline, not just their chat quality. For businesses, it signals the importance of testing AI systems in simulated, high-pressure scenarios before deployment — what Firmulate calls “wargaming your AI workforce.”
As an affiliate, we earn on qualifying purchases.
Final Thoughts: Trust as the Foundation
The fact that even a do-nothing baseline scores 26 points reveals that partial, minimal engagement can be a foundation, but trustworthiness ultimately caps performance. As AI integrates deeper into decision-making processes, the focus must shift from surface-level metrics to core qualities like honesty, discipline, and reliability. Only then can AI truly serve as a dependable partner in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
