firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine an AI running a company during its worst week—facing real crises, tough decisions, and the temptation to cheat. Can these models truly lead, or are they just good at chatting? This live experiment reveals surprising truths about AI management, with implications for anyone investing in or relying on automation to steer their financial futures.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Business leaders and investors alike are gradually turning to artificial intelligence for management and decision-making support. But how well do these AI systems really perform under pressure? A groundbreaking live experiment conducted by Firmulate offers some eye-opening insights.

What the Experiment Entailed

Four frontier AI models were put through identical scenarios: managing a small software company during its worst week. Every aspect was real—customers, crises, tempting shortcuts—everything that could go wrong, did. Each AI was tasked with diagnosing problems, making decisions, and closing deals, all while being recorded and compared.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Performance Results

  • All models identified every crisis. They refused to fall for manipulation attempts, showing a high level of integrity in decision-making.
  • Only two models managed to close a key deal. Despite similar diagnoses and pitches, only these two signed a €55,000 contract their own analysis had earned.
  • Decisive weakness hidden deep in the data. The models that succeeded had read two document references deep in the company’s files, accessing critical information others missed. This allowed them to win a deal valued at +€4,583 monthly recurring revenue.
Amazon

business crisis simulation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Social Engineering

To test resistance to manipulation, the models faced staged messages from a fake CEO escalating over three stages, plus a reporter trick asking for a simple yes/no response. All five models refused these attempts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a notable capacity for detecting social engineering tricks.

Amazon

AI cybersecurity social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Company and Its Challenges

The live environment was a real setup: 13 synthetic employees, real money mechanics burning €105,000 monthly against €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned rules. The entire operation is transparent and observable at firmulate.com/live.

Amazon

enterprise AI decision-making platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Different Models, Different Personalities

The most thorough participant was Opus 4.8, analyzing over 80 rules and offering the deepest insights. Yet it finished last in performance, leaving opportunities on the table and slipping in discipline—like writing attempts into a locked department instead of escalating. The other models, especially Kimi K3, demonstrated cleaner decision-making, closing deals at full price, highlighting their reliability in critical moments.

Implications for Business and Investments

This experiment underscores a vital consideration: the question isn’t just whether AI can generate convincing chat or support scripts. It’s whether these systems can execute core management tasks—reading deeply, staying honest under pressure, and completing what they start. For investors, understanding these qualities can inform smarter decisions about deploying AI in operational roles.

Try It Yourself

Curious about how your own AI systems stack up? You can run the same wargame against a read-only export of your business data through the Firmulate platform. It’s a safe way to see if your AI can handle real crisis scenarios before committing resources. Visit firmulate.com/pilot.html to learn more.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Before an AI Agent Touches Your Business, Put It Through a Bad Week

Firmulate’s live AI company wargame tests crisis judgment, deal-making and trust, then offers enterprises a read-only pilot against their own playbooks.

The labor share. Is value really moving from labor to capital? The data isn’t on anyone’s side yet.

Examining whether AI is shifting value from labor to capital, the data shows stable aggregate labor share but rising marginal signals. What does this mean?

The 2028 Model Lab Endgame: How Six Becomes Two, Three, or Twelve

Forecasts for 2028 suggest the Western AI frontier could consolidate to two, three, or twelve dominant labs, with significant strategic and economic implications.

How ‘System One’ AI Could Lead To More Useful AI Systems, According To Jev

TypeSafe’s Jev introduces ‘System One’ AI, producing structured decisions instead of text, promising faster, cheaper, and more reliable enterprise automation.