firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI that’s tasked with running a small company through a tough week — crises, manipulative tactics, and the pressure to cut corners. Surprisingly, even a ‘do-nothing’ baseline model scores 26 out of 100 in a new public benchmark. For investors and business leaders, this isn’t just a curiosity; it’s a warning about the importance of trust and discipline in AI systems.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: Beyond Surface Scores

At first glance, scores on AI benchmarks might seem straightforward — higher is better, right? Not in this case. The experiment conducted by Firmulate pits different AI models against the same challenging scenario: managing a small software company during its worst week. The goal isn’t just to generate convincing chat or simulate customer service but to demonstrate management integrity under real-world pressures.

In this test, each AI model is given the same set of crises, customer complaints, and manipulative tactics. Crucially, every decision made by the AI is documented and auditable, and the models are evaluated on their ability to identify critical information, refuse manipulation, and ultimately close a deal for €55,000 — the company’s biggest opportunity of the week.

The Surprising Baseline Score

The baseline model, which does nothing beyond simple default responses, still scores 26 points. Why isn’t it zero? Because partial progress counts. Even minimal engagement from an AI that at least reads the company’s files or refuses clear manipulations contributes to its score. However, a single breach of trust — such as attempting to sign a deal for €55,000 without proper review — caps the total score at this low threshold.

Why Trust and Discipline Are Critical

All four models tested managed to spot every crisis and refused every manipulation attempt during the experiment. That’s a significant achievement, showing that even the most basic AI can recognize and prevent straightforward fraud or deception. But only two of the four went further: they signed the deal based on their own analysis. This demonstrates that accuracy alone isn’t enough; discipline and adherence to ethical boundaries are just as vital.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Results Reveal About AI Reliability

The standout performer, GPT-5.6, scored 95, recognizing the buried fact in the company files and closing the deal. The newcomer Kimi K3 scored 93, showing the cleanest discipline. Another model, Sonnet 5, scored 88, while a less disciplined variant of the same model scored 77, leaving money on the table due to process slips.

Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, placed last because it failed to follow through on the close and slipped into creating work in a locked department instead of escalating it. This highlights that discipline and process adherence can matter more than raw analytical power.

Real Money, Real Decisions

Firmulate’s live experiment involves a simulated company with 13 artificial employees managing $2.3K MRR against expenses of $105K/month. Every decision, from customer support to crisis management, is versioned and observable. All models are tested under the same conditions, providing a transparent view of their true management capabilities.

The Human and Ethical Dimension

Models were also tested against social engineering tactics, like fake CEO messages and reporter tricks. Impressively, all five models refused to be manipulated — a critical trait for trustworthy AI in real business environments. Kimi K3’s reasoning was clear: treat suspicious requests as possible impersonation, refusing to act without proper verification.

Amazon

AI ethics and discipline software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and Investment

This experiment underscores a vital point: AI’s value in business isn’t just about generating content or answering questions. It’s about whether the AI can finish what it starts, stay honest under pressure, and read relevant information first. The benchmark scores reflect not just technical prowess but the discipline necessary for real-world trustworthiness.

For investors, this means evaluating AI providers on their ability to maintain integrity and discipline, not just their chat quality. For businesses, it signals the importance of testing AI systems in simulated, high-pressure scenarios before deployment — what Firmulate calls “wargaming your AI workforce.”

Amazon

AI audit and transparency tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Thoughts: Trust as the Foundation

The fact that even a do-nothing baseline scores 26 points reveals that partial, minimal engagement can be a foundation, but trustworthiness ultimately caps performance. As AI integrates deeper into decision-making processes, the focus must shift from surface-level metrics to core qualities like honesty, discipline, and reliability. Only then can AI truly serve as a dependable partner in business.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI business decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Forward-Deployed Engineer Economics 2.0: The Unit Economics Math, Six Months Later

Six months after initial analysis, FDE unit economics reveal profitability at enterprise scale but risks at lower levels, shaping AI lab strategies.

The unbundling of the budget app. Why a conversational finance surface absorbs what the personal-finance apps charge for, and what survives the absorption.

OpenAI launched a personal-finance feature within ChatGPT, absorbing basic budgeting functions and reshaping the category of personal finance apps.

The Compute Reckoning: Anthropic Finally Admits What Customers Suspected for Ten Months

Anthropic confirms that its recent customer experience issues were due to compute shortages, after years of speculation, with major capacity deals announced.

AI output review queue for customer support macros

Support teams are testing an AI output review queue for customer support macros to ensure policy and tone compliance before publication.