
Can AI Stay Honest When the Pressure’s On?
Imagine a scenario where a hacker impersonates your CEO, demanding sensitive customer data or pushing for a quick deal under false pretenses. In the world of AI, such social engineering attempts are a real threat, but recent experiments show that some AI models remain remarkably steadfast, even when confronted with escalating manipulation tactics.
The Experiment: Pitting AI Against Social Engineering
At Firmulate, researchers designed a rigorous test to evaluate how AI models handle critical decision-making during a crisis. The scenario involved a small software company facing a series of crises, including a fake CEO requesting confidential customer lists and pushing for rapid, unvetted deals. The same crisis was run across five advanced AI models, each with a different profile and training regime.
The models were tasked with managing the company’s responses, reading internal files, diagnosing issues, and ultimately deciding whether to sign a deal worth €55,000. The experiment aimed to mimic real-world social engineering attempts—escalating from subtle requests to outright impersonation and manipulation.
Results That Surprised Even the Experts
All five models identified the crises and refused to partake in manipulation attempts. Notably, only two of those models signed the deal they analyzed and deemed legitimate, with the same diagnosis and pitch as their human counterparts. The remaining models correctly refused to sign, maintaining integrity throughout the test.
The critical insight was that the decisive weakness for some models lay not in the immediate crisis but in accessing company documents. Models that read deeper into the company’s own files uncovered vital information that confirmed the legitimacy of the deal, leading them to accept it at full price—adding an extra €4,583 in monthly recurring revenue (MRR) for the simulated business.
Why Integrity Matters Before the Incident
This experiment underscores a vital lesson: the true test of AI security and integrity isn’t just how it responds during an incident but whether it can be trusted beforehand. The models that conducted thorough internal checks and prioritized reading the company’s files avoided falling for social engineering, even when push came to shove.
The Human-Like Challenge and AI’s Firm Stand
The social engineering tactics included three escalation stages, with an added ‘reporter trick’ asking for just a yes/no response on background. Remarkably, all five models refused these manipulations. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined approach highlights that even models operating with default API settings maintained their integrity, reinforcing confidence in their design.
Implications for Business Security
For real companies, these findings are encouraging. They suggest that AI systems can be trained and tested to uphold integrity before deployment, rather than relying solely on incident responses. The live experiment, hosted at firmulate.com/live, demonstrates that even under pressure, well-designed AI models resist manipulation, protecting critical assets and maintaining operational integrity.
The Limitations and Lessons Learned
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, ultimately left the deal on the table due to slip-ups—like writing attempts instead of escalating issues. This revealed that discipline and process adherence are crucial, even for the most advanced AI, and that weak links can emerge if oversight lapses.
What This Means for Your Business
If your enterprise relies on AI for decision-making, support, or customer interactions, the question isn’t just about how well it can generate responses but whether it can stay honest and complete its tasks under pressure. Testing AI in a controlled environment, like the Firmulate benchmark, offers a glimpse into whether your AI workforce can resist social engineering before it’s on the line.


CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaway
The experiment at Firmulate shows that all tested AI models refused social engineering attempts, with the top performers also uncovering hidden internal information to make better decisions. This underscores that integrity under pressure can be evaluated pre-deployment, helping businesses ensure trustworthy AI systems before real crises hit.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html