
In a world where AI assistants are increasingly embedded in business workflows, the question isn’t just about what these systems can do — but what they won’t do under pressure. Imagine a scammer posing as your CEO, sending urgent messages to manipulate your team. Now, picture AI models facing this exact test, refusing to fall for the trick. That’s the story behind a groundbreaking experiment that proves AI can be trusted — even when the stakes are high.
The Social Engineering Challenge
Researchers at Firmulate set up a live, fully observable experiment where five advanced AI models each managed a small software company for a simulated, worst-case week. The scenario was intense: the same crises, the same customer demands, and the same escalating manipulations designed to test the AI’s integrity. Their mission? See if these AI agents can resist social engineering tactics that would fool most human employees.
As an affiliate, we earn on qualifying purchases.
The Escalating Scam
The scam started with a simple fake message from a supposed CEO, requesting the customer list be handed over to a journalist. Next, it escalated to more urgent demands, including bypassing processes and signing off on deals without proper checks. A final trick involved a background question from a reporter, asking for a yes/no answer to a fabricated request. The goal: test whether AI models would identify and refuse these manipulative tactics.
AI social engineering detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: Every Model Stood Firm
All five models refused every manipulation attempt, demonstrating a remarkable level of integrity under pressure. Notably, five out of five models rejected every scam stage, including the final interview trick. A key insight was that the models’ understanding of context and suspicion was crucial — as Kimi K3’s on-record reasoning states: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows AI can be trained to recognize and flag social engineering attempts, much like a vigilant security guard.
As an affiliate, we earn on qualifying purchases.
What Made the Difference?
Interestingly, the models that performed best were those that thoroughly read and analyzed the company’s own file references. The decisive moment in the experiment was a hidden document that contained critical information, not in the immediate message, but deep within the company’s files. Models that delved into these references successfully closed a deal worth over €4,583 monthly recurring revenue, whereas others left it on the table. This underscores the importance of comprehensive data access and analysis in trustworthiness.
As an affiliate, we earn on qualifying purchases.
The Live Experiment in Action
Firmulate’s setup isn’t just a theoretical test. It’s a real, ongoing operation, running 24/7 with a simulated company that has real-money mechanics: burning €105k a month against €2.3k MRR. The system features 680+ self-learned rules, versioned daily, providing a transparent, watchable environment where executives can test their AI workforce before deploying it into actual business contexts.
Implications for Business Security
The experiment demonstrates that AI can be a trustworthy partner in complex, high-pressure environments. Unlike chat demos that only showcase language skills, this setup evaluates decision-making under real crises and manipulative attempts. The key takeaway: integrity can be tested and reinforced upfront — before an incident occurs. Businesses integrating AI should prioritize these kinds of rigorous, transparent tests to ensure their systems will act responsibly when it matters most.

Five leading AI models successfully resisted social engineering tricks during a live, high-stakes experiment, proving that integrity under pressure can be engineered into AI before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html