
Imagine a scenario where a fake CEO requests sensitive customer data, and your AI workforce refuses. In a world increasingly driven by automated decision-making, trust and integrity are no longer just human virtues—they are technological imperatives. Recent experiments with AI models reveal a promising trend: when tested against social engineering tricks, these systems not only recognize the threats but also refuse to act unethically.
Rethinking AI Security Through Live Wargaming
In a groundbreaking live experiment conducted by Firmulate, four advanced AI models were pitted against the toughest week a small software company might face. The goal? To see whether these AI systems could handle crises, temptations, and manipulation attempts without compromising integrity or operational goals. This was no theoretical test but a real-time simulation involving the same customers, crises, and unethical prompts across every model, with decisions fully documented and auditable.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Models Tested and Their Performance
- GPT-5.6-sol scored the highest at 95 points, successfully uncovering a hidden document reference that clinched a €55,000 deal, demonstrating full situational awareness and integrity.
- Kimi K3, the newcomer, scored 93 points, also closing the deal—its discipline in refusing unethical requests was the cleanest among all models.
- Sonnet 5 and Fable 5 scored 88 and 77, respectively, both managing to close deals but with minor slips in process discipline.
Crucially, all models identified every crisis and refused every attempt at manipulation, including staged social engineering escalations involving fake CEO messages and even a subtle reporter trick. The tests underscore that AI security isn’t just about detecting threats—it’s about ensuring systems adhere to ethical standards even under pressure.
The Hidden Weakness and Its Implications
The experiment revealed a subtle yet decisive weakness: models that could read deeply into internal files and references had a clear advantage. Those that managed to access and understand the company’s own documentation were able to close deals at full price, adding €4,583 MRR. Conversely, models with shallower access or discipline slipped in process discipline, leaving potential revenue on the table. This finding emphasizes the importance of comprehensive data access and compliance in AI decision-making—factors that can determine whether an AI system acts ethically and effectively.
Why This Matters for Business Leaders
As AI systems increasingly integrate into core business functions—be it CRM, customer support, or forecasting—the question isn’t whether they can generate convincing language or responses. It’s whether they can finish what they start, stay honest under pressure, and read critical internal information before acting. The live experiment by Firmulate shows that when tested rigorously, all five models refused unethical prompts, including staged attempts to manipulate them into illegal or unprofessional actions.
Insights from the Frontline of AI Integrity
The most thorough participant, Opus 4.8, demonstrated that even with the deepest analysis and rules learned, discipline can slip if the AI is not set up properly. In this case, the AI’s decision to leave a deal on the table stemmed from a reluctance to escalate suspicious requests, illustrating that internal process discipline is crucial.
Practical Steps for Organizations
To safeguard your business with AI, consider running your own internal ‘wargames’—simulations that test your AI’s response to ethical dilemmas and social engineering attempts before deployment. Firmulate’s platform offers a real, transparent environment where you can see how your models behave against real crises, with decisions fully auditable and decisions recorded. This proactive approach ensures your AI workforce maintains integrity and compliance when stakes are high.
Conclusion: Building Trust Before Incidents Happen
The key takeaway is that AI security isn’t just about responding to breaches—it’s about preventing them through rigorous testing and discipline. The live experiment demonstrates that with proper setup and deep understanding, AI models can be resilient against manipulation and uphold ethical standards. As firms look to automate more decision-making, this kind of pre-emptive testing will become essential to protect both reputation and revenue.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html