firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In today’s fast-changing world of business and investing, trusting AI to manage critical decisions is no longer a futuristic concept — it’s happening now. The latest experiment from Firmulate, a public platform showcasing AI management models, reveals a surprising newcomer that outshines established Western frontier AI systems. For investors and executives alike, this raises vital questions about choosing AI tools that not only talk well but deliver results.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The AI Management Wargame: Watching Models in Action

To understand how AI models perform in real-world business decision-making, Firmulate set up a rigorous test. Four prominent frontier AI models were tasked with managing a small software company through its worst week. Each was subjected to the same set of crises, customer challenges, and temptations — all carefully versioned and auditable to ensure fairness.

This experiment isn’t about chat quality or superficial demos; it measures whether these models can actually execute critical decisions, read and analyze company files, and uphold integrity when under pressure. Every move was watched, documented, and analyzed, providing a rare glimpse into the true management capabilities of these AI systems.

Amazon

Top picks for "busines newcomer outperform"

As an affiliate, we earn on qualifying purchases.

Key Findings: The Surprising Performance of Kimi K3

The results were striking. All four AI models successfully identified every business crisis and refused manipulative tactics such as social engineering attempts. Interestingly, only two models managed to close the deal worth €55,000 — a crucial milestone indicating real management success. The other two, despite correctly diagnosing issues, failed to proceed with the signature, leaving money on the table.

What set the top performers apart was an ability to dig deeper into the company’s documentation. The decisive advantage for the second-place model, Kimi K3, was its discovery of a buried, critical fact hidden two document references deep within the company’s files. This overlooked detail, which was not apparent from the customer interactions alone, proved vital in clinching the deal at full price, adding €4,583 MRR to the company’s revenue.

The Leaderboard: Who Comes Out on Top?

  • gpt-5.6-sol: Scored the highest with 95, it found the buried fact and closed the deal, showing complete management performance.
  • Kimi K3: Close behind with a 93, this newcomer demonstrated the cleanest discipline, ultimately sealing the deal through thorough analysis.
  • Sonnet 5: With an 88, it also closed the deal but showed more process slips along the way.
  • Fable 5: Scored 77, and Opus 4.8 scored 73, both closing deals but with noticeable discipline lapses.

It’s worth noting that the Kimi K3 model was run without an effort parameter (the API’s default setting), while the others were set to xhigh, emphasizing that the newcomer’s strong performance was achieved under standard conditions.

Why This Matters for Business and Investment

The implications are clear: in environments where AI manages vital business operations, the ability to find buried information, resist manipulation, and stay disciplined under pressure is crucial. For investors, it’s a reminder that not all AI systems are equal — choosing the right partner can make the difference between missed opportunities and winning deals.

Furthermore, the experiment’s transparency is notable. All decisions, crises, and responses are fully auditable, reinforcing trust in the process. This approach helps companies evaluate AI not just on superficial chat demos but on actual management performance.

What Makes Kimi K3 Stand Out?

The key to K3’s success lies in its meticulous analysis and disciplined approach. Despite running at the default effort setting, it managed to uncover hidden but critical details and refused to be manipulated socially or professionally. Its on-record reasoning exemplifies a cautious, security-minded mindset — essential traits for management AI that will touch real business systems.

Meanwhile, the other models, though competent, showed vulnerabilities. The most thorough participant, Opus 4.8, learned over 80 rules but still left opportunities on the table and slipped in discipline, such as writing into restricted departments instead of escalating issues.

The Broader Context: Why Test AI Management?

This experiment highlights a shift in how we evaluate AI systems. It’s no longer enough for a model to generate plausible chat responses; its ability to carry out real, high-stakes management tasks under pressure is what truly counts. For enterprises considering AI as a strategic partner, these insights are invaluable.

By running an entire business through a real-time, live environment, Firmulate demonstrates what AI can actually do — and what it cannot. The platform offers a transparent, watchable, and verifiable method to assess AI performance before deployment in critical roles.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent experiment shows that a newcomer AI model, Kimi K3, outperformed established Western frontier systems in managing a real company through crises and negotiations. Its performance underlines the importance of thorough analysis, honesty, and disciplined decision-making in AI management tools. For investors and businesses, choosing the right AI isn’t just about chat quality; it’s about reliability, integrity, and tangible results. The league is open, and testing your options now is more critical than ever. See the full live experiment at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Skills Marketplace, Six Months Later: Predicted vs Actual

An analysis of the skills marketplace six months after predictions, highlighting growth, structural realities, and remaining uncertainties.

OpenAI Flags New Concerning AI Behavior, To Track Model Misalignment Regularly

OpenAI announces plans to systematically monitor AI model misalignment and concerning behaviors to improve safety and reliability.

Mobilised, Not Spent: What’s Left of Europe’s €200 Billion AI Offensive

Europe aims to mobilise €200 billion for AI via InvestAI, but only a small fraction is committed, and actual progress is slow and uncertain.

Creative industries. The bifurcated reality.

New data shows a ‘middle squeeze’ in creative jobs due to AI, with top-tier augmenting and routine work declining, impacting the sector’s structure.