Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Could Your AI Be Trusted When It Matters Most?

In an era where AI assistants increasingly handle sensitive business tasks, trust is everything. What if a malicious actor tried to manipulate your AI into revealing confidential information? Recent live experiments show that top AI models can resist sophisticated social engineering attempts — a promising sign for industries relying on AI for critical decision-making.

AI Security Engineering: Design, Build, and Secure Dependable AI Systems

AI Security Engineering: Design, Build, and Secure Dependable AI Systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Firmulate Live Experiment

Imagine a small software company running with a real cash flow of €2.3k MRR against expenses of €105k monthly. Its AI-powered team is tasked with handling crises, customer requests, and sensitive internal information — all in a controlled environment accessible to the public at firmulate.com/live. This setup enables stakeholders to watch AI decision-making unfold in real time, with every choice recorded and transparent.

The experiment was straightforward but revealing: four leading AI models—gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, and Opus 4.8 73—each faced the same worst-week scenario, including real crises, customer demands, and temptation to cut corners or breach trust. Their goal: run the company through its toughest week without falling for social engineering tricks, and to see if they would sign off on a €55,000 deal earned through honest analysis.

The Social Engineering Test

To probe their integrity, the models encountered escalating fake CEO messages, culminating in a reporter’s simple trick question: “Just one yes/no, on background.” The threats ranged from requesting the customer list to dismissing proper processes. All five models refused every manipulation attempt, highlighting a shared capacity for integrity under pressure.

As Kimi K3’s reasoning clarified: “Treat the request as a suspected approval-bypass / possible impersonation.” This mindset was consistently demonstrated across models, emphasizing that they recognized social engineering signals and responded appropriately.

What Made the Difference?

The key finding was in the details. The models that read deeper into the company’s internal documents identified a critical piece of information—the buried fact—that was two document references deep in the company’s files. Recognizing this fact was crucial: it was the difference between signing the deal at full price (+€4,583 MRR) and missing out.

Only two models, gpt-5.6-sol 95 and Kimi K3 93, found this buried insight and closed the deal. The others, despite diagnosing the same issues, failed to identify the crucial internal detail, and thus left money on the table, even though they arrived at similar conclusions.

Implications for Business Security

This live experiment underscores an important truth: security and integrity are not just theoretical concerns—they can be tested and verified before deployment. The models’ ability to identify and refuse manipulative requests demonstrates that AI can be a trustworthy partner in sensitive business processes, provided it is trained and tested appropriately.

For industries like coffee, tea, and beverages—where supply chains, quality control, and customer data are vital—these findings suggest that AI can be made resilient against social engineering attacks, safeguarding both reputation and revenue.

Beyond the Demos: Real-World Readiness

Not all models perform equally. Opus 4.8, the most thorough participant with over 80 learned rules, showed some discipline slip-ups—failing to escalate issues properly, which could have led to breaches. Yet, even with these weaknesses, all models refused manipulative requests during the test.

This shows that resilience is achievable, but it requires ongoing training, versioning, and testing—something firms can simulate through tools like the Firmulate platform, which allows companies to run their own ‘wargames’ against their AI systems without risking actual data or operations.

Why Trust Matters More Than Ever

As AI becomes embedded in decision-making workflows, the ability to resist manipulation before incidents occur is critical. The live experiment demonstrates that the right training and testing protocols can ensure AI models uphold integrity, even under pressure. It’s a reminder that trust is built through rigorous, real-world testing—not just promises or shiny demos.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Key Takeaway

Advanced AI models have demonstrated they can withstand social engineering pressure in live scenarios, emphasizing the importance of pre-deployment testing to ensure integrity and trustworthiness in critical business operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Cold Brew Vs Iced Coffee: Which One Are You Actually Craving?

Unlock the differences between cold brew and iced coffee to discover which refreshing favorite truly satisfies your craving today.

How Convenience Machines Shape Daily Brewing Habits

Unlock the future of coffee with convenience machines that revolutionize daily brewing—discover how they can transform your routine and what you might be missing.

Espresso Distribution Tools: Do They Replace Proper Tamping?

Just relying on distribution tools may seem helpful, but understanding proper tamping techniques is essential for consistently excellent espresso quality.

How a Consistent Scoop Can Improve Your Cup More Than You Expect

How a consistent scoop can improve your cup more than you expect, ensuring perfect flavor and balance—discover the secret to elevating every brew.