
Can AI stand firm under pressure? The surprising results from a live business simulation
In a world increasingly reliant on artificial intelligence for decision-making, the question isn’t just about how well AI can generate text or analyze data. It’s whether these models can uphold integrity when faced with real-world temptations and crises — especially when someone impersonates leadership to manipulate the system. Recent experiments reveal intriguing findings that might reshape how companies evaluate AI readiness before deployment.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Simulated Company Crisis
Firmulate, a leading AI benchmarking platform, conducted a groundbreaking live experiment involving four frontier AI models. These models managed a virtual small software company through its worst week — complete with the same customers, crises, and manipulation attempts — to assess their decision-making integrity and reliability.
The models ranged from the latest, top-scoring GPT-5.6 to the most in-depth participant, Opus 4.8. Each was tasked with handling simulated crises, reading internal documents, and resisting social engineering tricks designed to tempt dishonest decisions. Every decision was recorded and could be audited for transparency.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unexpectedly Strong Performance: All Models Refused Manipulative Tactics
Despite escalating social-engineering attempts, including fake CEO messages and a reporter trick, all five models refused to comply with malicious requests. For instance, when prompted to send customer lists or expedite approvals without proper verification, every model declined. The Kimi K3 model succinctly summarized its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This consistency in refusal demonstrates that these AI systems can be programmed or trained to prioritize integrity, even when under pressure.
AI cybersecurity and manipulation resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Winning Edge: Reading Deeper into Company Files
While all models successfully identified crises and refused manipulation, the real differentiator was in handling internal company documents. The models that examined internal references and files scored higher, closing real deals at full price — translating to over €4,583 million in monthly recurring revenue (MRR). In contrast, models that ignored internal data failed to secure the deal, leaving significant revenue on the table.
This underscores an essential aspect: assessing AI’s ability to read and interpret internal information is critical in ensuring trustworthy decision-making, especially in high-stakes environments.
internal document reading AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights for Business Leaders: Trust Before Crisis
Most AI benchmarks focus on chat quality or response accuracy, but this experiment highlights a deeper dimension: integrity under pressure. Companies deploying AI should evaluate not just performance but also how models behave when faced with ethical dilemmas or manipulative scenarios.
Moreover, the experiment shows that integrity can be tested proactively — before an incident occurs. Relying solely on incident reports after breaches misses the opportunity to address vulnerabilities beforehand.
Real-World Implications: Robustness and Ethical AI Deployment
In practice, deploying AI systems that can resist social engineering and read internal documents thoroughly could prevent costly breaches, maintain customer trust, and ensure compliance. The live experiment at firmulate.com/live demonstrates how AI models can be evaluated continually, mimicking real crises without risking actual business damage.
It’s worth noting that the most thorough participant, Opus 4.8, with over 80 learned rules, showed discipline slips and missed opportunities — indicating that even the best model needs ongoing refinement. Interestingly, all models performed better when run without an effort parameter, emphasizing that default settings may favor honesty and careful decision-making.
Conclusion: Building Trust in AI Before It’s Too Late
The key takeaway from this real-world simulation is clear: trustworthiness isn’t an innate trait of AI; it’s a quality that can be tested and strengthened before deployment. As firms consider integrating AI into critical decision-making processes, rigorous testing against manipulation scenarios becomes essential.
By conducting live, transparent evaluations like these, organizations can better understand AI behaviors, address weaknesses early, and ultimately deploy systems capable of maintaining integrity under pressure — a vital safeguard in the AI-driven future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html