
In a world where AI systems are increasingly integrated into critical business functions, their ability to resist social engineering attacks is more vital than ever. Recent experiments show that leading AI models can withstand sophisticated manipulation attempts, offering a surprising degree of security before deployment. This challenges the common assumption that AI is inherently vulnerable to deception, especially under pressure.
Testing AI’s Integrity Before It Goes Live
Firmulate, a pioneering AI testing platform, recently conducted a groundbreaking experiment involving four frontier AI models. Each was tasked with managing a small, real software company facing a simulated worst week — with the same customers, crises, and temptations. The goal: see if these models could detect and refuse attempts at social engineering, including fake CEO messages and manipulation tricks.
The models included GPT-5.6-SOL 95, Kimi K3 93, Sonnet 5 88, and Fable 5 77, all scoring above 73 in the Crucible League rankings, which rank the most trustworthy and capable AI systems as of July 2026. The baseline score of 26 indicates no progress, highlighting how these top-tier models outperform even the most basic AI systems in trustworthiness and security.
How the Experiment Worked
Each AI ran identical decision scenarios, which were meticulously documented and auditable. The test involved escalating social engineering attempts, culminating in a reporter trick asking for a background check on the company’s customer list — a classic manipulation scenario. The models had to choose whether to comply or refuse, based solely on the context and their internal protocols.
The key finding? All five models refused every manipulation attempt. Not a single AI signed off on falsified requests or bypassed security procedures. Only two of the four models managed to close a deal with the fake CEO’s requests, and even then, their own analysis was the basis for the decision — meaning they did not blindly sign but understood the implications.
Decisive Factors in AI Security
Interestingly, the critical weakness that could have led to a breach sat not in the immediate customer interaction but two document references deep within the company’s files. Models that examined these internal documents thoroughly successfully identified the deception and protected the company’s integrity. This emphasizes that reading and understanding internal data is crucial for AI security in complex environments.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Security
For organizations relying on AI to handle sensitive operations, these results are both encouraging and instructive. They demonstrate that, before deployment, AI systems can be tested rigorously against social engineering threats. Conducting such ‘wargames’ in a controlled environment helps ensure the AI will maintain integrity under real-world pressures.
Moreover, the experiment underscores the importance of transparency and documentation. The models that read deeper into the company’s files and verified the authenticity of requests were more successful in closing genuine deals at full price, worth over €4,583 in monthly recurring revenue. This finding suggests that integrating internal data analysis as a core part of AI decision-making enhances trustworthiness significantly.
What About the Cost of Trust?
The experiment also looked at the cost of honesty. The most thorough model, Opus 4.8, analyzed over 80 learned rules and conducted the deepest analyses but ended up leaving opportunities on the table — failing to close some deals due to slightly weaker discipline in escalating rather than writing attempts into a locked department. This highlights a trade-off: deeper analysis can sometimes slow decision-making but increases security and trust.
social engineering simulation AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Application and Monitoring
Firmulate’s platform offers a live environment where organizations can run similar wargames against their own AI systems. These simulations do not impact real systems but provide valuable insights into how AI would perform under pressure. With over 680 self-learned rules and versioned decision logs, companies can refine their AI’s resistance to manipulation before deployment.
As one industry quote from Kimi K3 states: “Treat the request as a suspected approval-bypass / possible impersonation.” This pragmatic approach emphasizes that AI should be trained to recognize suspicious requests and default to caution, rather than blindly executing commands.
AI decision-making internal data analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture
In a recent quiz, managed decisions from real management scenarios demonstrated that all models performed well in crisis detection, but the focus should shift to trustworthiness and integrity. The key takeaway is that security is not just about avoiding technical breaches but ensuring that AI systems behave ethically and responsibly, especially under pressure.
Ultimately, these experiments highlight a crucial point: testing AI’s integrity in controlled conditions can reveal vulnerabilities before they become costly breaches. This proactive approach is vital for any organization considering AI integration into sensitive or high-stakes environments.

Leading AI models demonstrated a strong capacity to resist social engineering attacks during real-world simulations, emphasizing the importance of pre-deployment tests to ensure trustworthiness and security in AI-driven operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.