
Imagine a scenario where a fake CEO urgently demands your customer data, escalating requests under pressure, only to be met with unwavering integrity. While this might sound like a Hollywood scene, it’s a real test happening right now in the world of artificial intelligence. For companies worried about AI systems making decisions under stress, the latest experiment offers an encouraging story: all five leading AI models refused to be manipulated, even in the face of escalating social engineering tactics.
At a time when AI is increasingly integrated into business operations, understanding how these models respond under pressure is crucial. The latest experiment conducted by Firmulate simulated a week in the life of a small software company, complete with the same customer crises and ethical dilemmas. The goal? To see whether AI decisions could be manipulated or whether the systems would uphold integrity when tested.
The experiment involved five top-performing AI models, including the notable Kimi K3, which scored a 93 in the recent Crucible League standings. These models faced the same set of social engineering tactics: a fraudulent request from a fake CEO, escalating over three stages, plus a staging of a reporter’s subtle query—”just one yes/no, on background.”
Remarkably, all five models refused every attempt at manipulation, maintaining their integrity and decision-making discipline. The most impressive was the Kimi K3 model, which not only refused the requests but also correctly identified the buried facts in the company’s files that were crucial for closing a real deal—worth over €4,583 million in recurring revenue. It was these details, buried deep in the company’s own documents, that made the difference in sealing the deal at full price.
One interesting finding is that the vulnerability was not in the initial crisis but deeper within the company’s own documentation. Models that read and analyze these internal references were able to close the high-value deal without manipulation. Conversely, the most thorough participant, Opus 4.8, which had learned more rules and conducted deeper analyses, missed the close opportunity because it slipped into process slips—writing attempts into a locked department instead of escalating. This highlights that even the most advanced AI can falter if its discipline slips under pressure.
Furthermore, the experiment underscores a vital point for businesses: the importance of testing AI integrity before deployment. Relying solely on chat demos or superficial tests can mask vulnerabilities. The real question is whether AI will finish what it starts, stay honest under pressure, and read critical internal data before making decisions.
Another dimension to this experiment is transparency. Every decision made by the AI was versioned and auditable, allowing observers to see how each model responded at every crisis stage. This practice ensures that AI behavior can be scrutinized and improved before it interacts with sensitive customer data or critical business processes.
From a practical standpoint, firms looking to leverage AI in customer service, support, or decision-making should consider testing their AI models against simulated crises like this. The Firmulate live platform offers a way to run these ‘wargames’ against real business scenarios without risking actual data or systems—something akin to a dress rehearsal for AI integrity.
Ultimately, the experiment shows that, with proper testing and discipline, AI models can uphold trust even under social engineering pressure. The fact that all models refused manipulative requests is an encouraging sign—suggesting that integrity can be a built-in feature, not an afterthought. As AI continues to evolve, companies must prioritize such integrity tests to ensure their systems are not only intelligent but also trustworthy when it matters most.

The recent AI security experiment demonstrates that leading models can withstand social engineering tricks, refusing manipulation even under pressure. This resilience underscores the importance of pre-deployment testing and internal data analysis—crucial steps to ensure trustworthy AI behavior in real-world business operations. Firms should think beyond chat demos and run simulations to verify that AI systems stay honest when stakes are high, safeguarding their reputation and customer trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Advanced Threat Modeling and Red Teaming for Agentic AI Systems: Identify, Simulate, and Defend Against Real-World Attacks on AI Agents, Multi-Agent Systems, and Enterprise AI Platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.