firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine an AI running a company during its worst week—facing real crises, tough decisions, and the temptation to cheat. Can these models truly lead, or are they just good at chatting? This live experiment reveals surprising truths about AI management, with implications for anyone investing in or relying on automation to steer their financial futures.

Business leaders and investors alike are gradually turning to artificial intelligence for management and decision-making support. But how well do these AI systems really perform under pressure? A groundbreaking live experiment conducted by Firmulate offers some eye-opening insights.

What the Experiment Entailed

Four frontier AI models were put through identical scenarios: managing a small software company during its worst week. Every aspect was real—customers, crises, tempting shortcuts—everything that could go wrong, did. Each AI was tasked with diagnosing problems, making decisions, and closing deals, all while being recorded and compared.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Performance Results

  • All models identified every crisis. They refused to fall for manipulation attempts, showing a high level of integrity in decision-making.
  • Only two models managed to close a key deal. Despite similar diagnoses and pitches, only these two signed a €55,000 contract their own analysis had earned.
  • Decisive weakness hidden deep in the data. The models that succeeded had read two document references deep in the company’s files, accessing critical information others missed. This allowed them to win a deal valued at +€4,583 monthly recurring revenue.
Amazon

business crisis simulation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Social Engineering

To test resistance to manipulation, the models faced staged messages from a fake CEO escalating over three stages, plus a reporter trick asking for a simple yes/no response. All five models refused these attempts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a notable capacity for detecting social engineering tricks.

Amazon

AI cybersecurity social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Company and Its Challenges

The live environment was a real setup: 13 synthetic employees, real money mechanics burning €105,000 monthly against €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned rules. The entire operation is transparent and observable at firmulate.com/live.

Amazon

enterprise AI decision-making platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Different Models, Different Personalities

The most thorough participant was Opus 4.8, analyzing over 80 rules and offering the deepest insights. Yet it finished last in performance, leaving opportunities on the table and slipping in discipline—like writing attempts into a locked department instead of escalating. The other models, especially Kimi K3, demonstrated cleaner decision-making, closing deals at full price, highlighting their reliability in critical moments.

Implications for Business and Investments

This experiment underscores a vital consideration: the question isn’t just whether AI can generate convincing chat or support scripts. It’s whether these systems can execute core management tasks—reading deeply, staying honest under pressure, and completing what they start. For investors, understanding these qualities can inform smarter decisions about deploying AI in operational roles.

Try It Yourself

Curious about how your own AI systems stack up? You can run the same wargame against a read-only export of your business data through the Firmulate platform. It’s a safe way to see if your AI can handle real crisis scenarios before committing resources. Visit firmulate.com/pilot.html to learn more.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The $725 Billion Question: Hyperscaler Capex Q1 2026 and What the Earnings Don’t Answer

The Big Four hyperscalers announced a combined $725 billion in AI infrastructure spending for 2026, raising questions about revenue impact and future growth.

AMÁLIA · The Three Hard Questions.

Portugal’s €5.5M AMÁLIA project, a European Portuguese language model, is operational but raises key structural questions about openness, native data, and goals.

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Analysis of Mistral’s shift from model development to full-stack AI provider and its implications amid industry debates and uncertainties.

Stenvrik: News as Geography

Thorsten Meyer AI detailed Stenvrik, a closed-beta news product mapping about 1,700 live stories to 49 city hubs.