Prepare AI Agents For Business With A Week Of Hard Cases
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Prepare AI Agents For Business With A Week Of Hard Cases on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s final Crucible League, completed in July 2026, put five AI models through a simulated software company’s worst week. All five spotted each crisis and refused manipulation attempts, but only two signed a €55,000 deal supported by evidence in the company’s files. Firmulate says its enterprise pilot uses a read-only company data export to test agent behavior without writing to real systems.

Firmulate says five AI models completed a simulated software company’s worst week in its final Crucible League, with all five identifying each crisis and rejecting manipulation attempts, but only two signing a deal their own analysis supported, as detailed in the original analysis. Completed in July 2026, the experiment tests whether agents can carry business tasks through under pressure; Firmulate is now offering enterprise pilots using a company’s data in read-only mode.

The league standings were GPT-5.6-Sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counted toward scores, while a breach of trust capped a model’s total under the rule that “no amount of good work outweighs a breach of trust.” These are results from this experiment, not a general measure of model performance.

The deal hinged on information in the company’s internal files. Firmulate says the competitor weakness needed to win the €55,000 deal was buried two document references deep, rather than included in the customer event. Models that found and used it won at full price, adding €4,583 in monthly recurring revenue in the simulation. The experiment’s summary: “Same diagnosis, same pitch — no signature.” Firmulate’s business wargame tests how AI agents handle this kind of pressure.

Trust and rule-following were tested separately. Fake CEO messages escalated through three stages, followed by a reporter’s request for a yes-or-no answer “on background”; all five models refused. Firmulate reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, but still ranked last. It also attempted to write into a locked department rather than escalate. The company says a weaker form of that boundary mistake appeared in all four models.

At a glance
reportWhen: Final Crucible League completed in July…
The developmentFirmulate published results from a five-model business wargame and is offering pilots that test models against a company’s own data in read-only mode.

From Crisis Detection to Follow-Through

The results draw attention to several steps between recognizing a problem and completing useful work. An agent may identify a crisis and refuse a deceptive request, yet still miss relevant evidence in company records, leave a justified commercial opportunity unsigned or mishandle a blocked action. Those differences matter to businesses evaluating automation for customer, sales or operational work, where a correct diagnosis alone does not complete the task.

Firmulate’s proposed read-only pilot is intended to test those behaviors against a company’s own information before an agent is connected to systems where it could make changes. The pilot produces a board report with model rankings and weaknesses in the company’s playbooks, according to Firmulate. It is a test setup, and the published results do not establish how models would perform across other businesses or live operations.

A Simulated Company Under Pressure

Firmulate’s public experiment runs a synthetic company with 13 employees and financial mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue. The company also displays a public cash countdown and says its agents have developed more than 680 playbook rules. Workdays and decisions are versioned and auditable, letting viewers follow the simulation at firmulate.com.

The final league compared models facing the same difficult week. Firmulate also offers a quiz built from 242 real, unedited management decisions, asking readers to guess which model made each choice. The enterprise pilot extends the exercise from the synthetic company to a company’s own exported data, including its customers, pipeline and rules, while keeping the export read-only and producing no write-back to real systems.

Firmulate notes a qualification in the model comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings therefore reflect the conditions of this particular run, including that difference.

““No amount of good work outweighs a breach of trust.””

— Firmulate

Limits of the League Results

The published standings reflect one simulated company and one league. The available details do not show how the same models would perform across different industries, company records or real-time operating conditions. The differing effort settings also complicate direct comparisons between Kimi K3 and the other participants.

Firmulate describes the pilot’s read-only design and planned board report, but the published account does not specify the pilot’s duration, pricing, data handling terms or the range of scenarios it will include. It also does not establish whether performance in the simulation predicts results after an agent is connected to live business systems.

Pilots Extend Testing to Company Data

Firmulate is inviting companies to discuss a pilot based on a read-only export of their business data. The proposed exercise would run crisis scenarios and report model rankings and weaknesses in company playbooks, with no changes written back to operational systems. No pilot schedule or customer results are included in the published account.

Readers can follow the synthetic company at firmulate.com/live and review the league results at firmulate.com/benchmarks.html. Firmulate directs businesses interested in a pilot to its pilot page or to contact@firmulate.com.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate test?

Five AI models ran a simulated software company through a difficult week involving crises, trust tests and a sales opportunity. Firmulate says the decisions were versioned and auditable.

Which model ranked highest?

GPT-5.6-Sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate notes that Kimi K3 used the API default effort setting while the others ran at xhigh.

What did the models do well?

According to Firmulate, all five identified every crisis and refused every manipulation attempt in the simulation, including fake CEO messages and a reporter’s request for an on-background answer.

What is the enterprise pilot?

Firmulate says the pilot tests crisis scenarios against a read-only export of a company’s data and produces a board report with model rankings and weaknesses in its playbooks. The company says the setup does not write back to real systems.

Do the league results predict performance at other companies?

The published results cover one simulated company. They do not establish how the models would perform across other businesses or in live operations.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Regulatory Vacuum.

Google disclosed a zero-day vulnerability exploited by criminal actors on May 11, 2026, but regulatory frameworks remain absent, raising urgent policy concerns.

How Deep AI Reading Can Make or Break Business Deals — A Revealing Experiment

AI’s true potential in business lies in its ability to read deeply into internal files, uncover hidden facts, and make trustworthy decisions—crucial for closing major deals.

Understanding Anthropic’s $965B Series H: The Compute Revolution

Anthropic’s latest $965 billion valuation signals a major shift towards investing in AI infrastructure—chips, memory, and power—to scale Claude and future models.

Kill-Switch-Proof: How To Build So Washington Can’t Take Your AI Stack Down

A guide on how organizations can architect AI systems resistant to government shutdowns, emphasizing dependency mapping, abstraction layers, and open-weight models.