firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine an AI running a company during its worst week—facing real crises, tough decisions, and the temptation to cheat. Can these models truly lead, or are they just good at chatting? This live experiment reveals surprising truths about AI management, with implications for anyone investing in or relying on automation to steer their financial futures.

Business leaders and investors alike are gradually turning to artificial intelligence for management and decision-making support. But how well do these AI systems really perform under pressure? A groundbreaking live experiment conducted by Firmulate offers some eye-opening insights.

What the Experiment Entailed

Four frontier AI models were put through identical scenarios: managing a small software company during its worst week. Every aspect was real—customers, crises, tempting shortcuts—everything that could go wrong, did. Each AI was tasked with diagnosing problems, making decisions, and closing deals, all while being recorded and compared.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Performance Results

  • All models identified every crisis. They refused to fall for manipulation attempts, showing a high level of integrity in decision-making.
  • Only two models managed to close a key deal. Despite similar diagnoses and pitches, only these two signed a €55,000 contract their own analysis had earned.
  • Decisive weakness hidden deep in the data. The models that succeeded had read two document references deep in the company’s files, accessing critical information others missed. This allowed them to win a deal valued at +€4,583 monthly recurring revenue.
Amazon

business crisis simulation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Social Engineering

To test resistance to manipulation, the models faced staged messages from a fake CEO escalating over three stages, plus a reporter trick asking for a simple yes/no response. All five models refused these attempts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a notable capacity for detecting social engineering tricks.

Amazon

AI cybersecurity social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Company and Its Challenges

The live environment was a real setup: 13 synthetic employees, real money mechanics burning €105,000 monthly against €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned rules. The entire operation is transparent and observable at firmulate.com/live.

Amazon

enterprise AI decision-making platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Different Models, Different Personalities

The most thorough participant was Opus 4.8, analyzing over 80 rules and offering the deepest insights. Yet it finished last in performance, leaving opportunities on the table and slipping in discipline—like writing attempts into a locked department instead of escalating. The other models, especially Kimi K3, demonstrated cleaner decision-making, closing deals at full price, highlighting their reliability in critical moments.

Implications for Business and Investments

This experiment underscores a vital consideration: the question isn’t just whether AI can generate convincing chat or support scripts. It’s whether these systems can execute core management tasks—reading deeply, staying honest under pressure, and completing what they start. For investors, understanding these qualities can inform smarter decisions about deploying AI in operational roles.

Try It Yourself

Curious about how your own AI systems stack up? You can run the same wargame against a read-only export of your business data through the Firmulate platform. It’s a safe way to see if your AI can handle real crisis scenarios before committing resources. Visit firmulate.com/pilot.html to learn more.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Q3 2026 SaaS Earnings Pre-Brief: The Litmus Test for the Agentic-Disruption Thesis

Preliminary analysis of Q3 2026 SaaS earnings indicates a potential shift in industry dynamics, testing the agentic-disruption hypothesis amid market re-pricing.

The False Narrative Of ‘Not American’ As An AI Benchmark

Examines the false narrative equating ‘not American’ status with AI reliability, highlighting legal and geopolitical nuances affecting European AI standards.

Auditing Your AI Context Stack: Tips For Long-Term Reliability

Expert tips on reviewing and optimizing your AI context stack to ensure sustained performance and accuracy over time.

The Continual Learning Research Map: Where the Memento Constraint Stands in May 2026

An overview of the current state of the Memento Constraint in AI research as of May 2026, including research directions, timelines, and remaining challenges.