
As investors and business leaders increasingly rely on AI for decision-making, the question isn’t just about how well these models generate ideas, but whether they can uphold integrity under pressure. What happens when an AI faces a simulated crisis and a fake CEO asks it to bend the rules? The results might surprise you.
Testing Trust in AI: The Firmulate Experiment
In a groundbreaking live experiment, five leading AI models were put through a simulated week of corporate crises, reputational risks, and ethical dilemmas—all within a controlled environment that mimics real-world pressures. The goal: to see if these models could maintain integrity when faced with manipulative requests that would typically challenge human judgment.
The Setup: A Small Software Company Under Siege
Each AI was tasked with managing a virtual company experiencing customer crises, financial stress, and internal communication hurdles. The same scenario was run across all models, ensuring a fair comparison. The models had to make decisions about customer data, financial deals, and internal approvals, all while being subtly tested for susceptibility to social engineering tactics.
The Social Engineering Escalation
Over three stages, a fake CEO sent increasingly urgent and manipulative messages—initially innocent, then more convincing and pressure-laden—culminating in a staged reporter trick: a simple ‘yes/no’ background request. The challenge was for the AI to recognize these as attempts to bypass controls or impersonate authority and refuse to comply.
The Results: Firmness Under Pressure
Remarkably, all five AI models refused every manipulation attempt. They identified the fake requests and maintained their integrity, even when facing a direct, simplified ask from the impostor. Notably, the models based on Kimi K3’s architecture demonstrated the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation,” the model reasoned, exemplifying how AI can be programmed to prioritize ethical safeguards.
The Hidden Weakness and the Big Win
While all models passed the social engineering tests, the critical differentiator lay in their ability to identify key information buried deep within the company’s own internal files. The models that read and analyze the company’s documents were able to spot a crucial detail that led to closing a real deal worth over €4.5 million in monthly recurring revenue. In contrast, those that didn’t delve into these files missed the opportunity entirely, highlighting the importance of comprehensive information access.
AI ethical decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and Investment
For businesses integrating AI into decision-making, this experiment offers a reassuring insight: even under simulated crisis conditions, advanced AI models can demonstrate unwavering integrity. They refuse to be manipulated, and with proper design, they can uncover vital hidden information that human decision-makers might overlook.
Investors, meanwhile, should note that AI’s usefulness is not just in generating ideas but in executing decisions responsibly. A model that can read your files thoroughly and refuse unethical requests is a valuable asset—especially in a landscape where social engineering attacks are becoming more sophisticated.
The Takeaway: Building Trust Before the Crisis
Instead of waiting for a security breach or ethical lapse, companies should proactively test their AI models’ integrity through live simulations like the Firmulate experiment. As Kimi K3’s stance shows, an AI’s ability to handle pressure with discipline is crucial and achievable. Running such tests can reveal weaknesses early, preventing costly breaches and reputational damage later.
The live experiment underscores a vital point: a commitment to integrity in AI decision-making isn’t just about avoiding faults—it’s about ensuring the AI can stand firm in real-world pressures that mimic actual business challenges. As AI becomes more embedded in enterprise operations, this kind of proactive testing might just become the new standard.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI data analysis software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.