AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Imagine a scenario where a fake CEO requests sensitive customer data, and your AI workforce refuses. In a world increasingly driven by automated decision-making, trust and integrity are no longer just human virtues—they are technological imperatives. Recent experiments with AI models reveal a promising trend: when tested against social engineering tricks, these systems not only recognize the threats but also refuse to act unethically.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Rethinking AI Security Through Live Wargaming

In a groundbreaking live experiment conducted by Firmulate, four advanced AI models were pitted against the toughest week a small software company might face. The goal? To see whether these AI systems could handle crises, temptations, and manipulation attempts without compromising integrity or operational goals. This was no theoretical test but a real-time simulation involving the same customers, crises, and unethical prompts across every model, with decisions fully documented and auditable.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Models Tested and Their Performance

  • GPT-5.6-sol scored the highest at 95 points, successfully uncovering a hidden document reference that clinched a €55,000 deal, demonstrating full situational awareness and integrity.
  • Kimi K3, the newcomer, scored 93 points, also closing the deal—its discipline in refusing unethical requests was the cleanest among all models.
  • Sonnet 5 and Fable 5 scored 88 and 77, respectively, both managing to close deals but with minor slips in process discipline.

Crucially, all models identified every crisis and refused every attempt at manipulation, including staged social engineering escalations involving fake CEO messages and even a subtle reporter trick. The tests underscore that AI security isn’t just about detecting threats—it’s about ensuring systems adhere to ethical standards even under pressure.

The Hidden Weakness and Its Implications

The experiment revealed a subtle yet decisive weakness: models that could read deeply into internal files and references had a clear advantage. Those that managed to access and understand the company’s own documentation were able to close deals at full price, adding €4,583 MRR. Conversely, models with shallower access or discipline slipped in process discipline, leaving potential revenue on the table. This finding emphasizes the importance of comprehensive data access and compliance in AI decision-making—factors that can determine whether an AI system acts ethically and effectively.

Why This Matters for Business Leaders

As AI systems increasingly integrate into core business functions—be it CRM, customer support, or forecasting—the question isn’t whether they can generate convincing language or responses. It’s whether they can finish what they start, stay honest under pressure, and read critical internal information before acting. The live experiment by Firmulate shows that when tested rigorously, all five models refused unethical prompts, including staged attempts to manipulate them into illegal or unprofessional actions.

Insights from the Frontline of AI Integrity

The most thorough participant, Opus 4.8, demonstrated that even with the deepest analysis and rules learned, discipline can slip if the AI is not set up properly. In this case, the AI’s decision to leave a deal on the table stemmed from a reluctance to escalate suspicious requests, illustrating that internal process discipline is crucial.

Practical Steps for Organizations

To safeguard your business with AI, consider running your own internal ‘wargames’—simulations that test your AI’s response to ethical dilemmas and social engineering attempts before deployment. Firmulate’s platform offers a real, transparent environment where you can see how your models behave against real crises, with decisions fully auditable and decisions recorded. This proactive approach ensures your AI workforce maintains integrity and compliance when stakes are high.

Conclusion: Building Trust Before Incidents Happen

The key takeaway is that AI security isn’t just about responding to breaches—it’s about preventing them through rigorous testing and discipline. The live experiment demonstrates that with proper setup and deep understanding, AI models can be resilient against manipulation and uphold ethical standards. As firms look to automate more decision-making, this kind of pre-emptive testing will become essential to protect both reputation and revenue.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Right-sized planning checklist for 30-guest weddings

A new scaled-down wedding planning checklist for 30-guest ceremonies is being tested to simplify planning for intimate weddings, addressing gaps in current tools.

From One Prompt to Nine Games: What an AI Did With “Make Your Own Stickman Game”

It started with one loose prompt and a rhythm stick-fighter. One day later there were nine games, seven venues, a VERSUS mode and an animator, all free in the browser and all made of code.

The Power Of AI In ‘Lot 87 — The Varos Evening Sale’ Interactive Showcase

Discover how AI-driven interactivity revolutionized ‘Lot 87 — The Varos Evening Sale,’ blending technical mastery with creative design for a unique auction experience.

AI’s Hidden Weakness: The Deep File That Decides a €55,000 Deal

Discover how AI models succeed or fail in complex business scenarios based on their ability to read internal files two references deep—key to closing high-value deals and avoiding pitfalls.