
Imagine you’re lifting weights at the gym, aiming for a new personal best. But instead of pushing yourself, you do nothing — yet somehow, you still get a score of 26 out of 100. Sounds odd, right? In the world of AI benchmarking, that’s exactly how trust and performance are measured. It’s an honest, if counterintuitive, approach that highlights what AI models truly deliver — and what they can’t.
Get workout gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Reality of AI Performance Benchmarks
In evaluating artificial intelligence for business management, the biggest revelation isn’t just how well models perform when everything goes smoothly; it’s how they handle crises, manipulations, and ethical dilemmas. The latest experiment from firmulate.com offers a fascinating look. Four leading AI models were tested by running a simulated small software company through its worst week, complete with customer crises, internal temptations, and ethical tests.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the ‘Do-Nothing’ Baseline Isn’t Zero
You might expect a baseline score of zero for an AI that does nothing. But surprisingly, the do-nothing approach scored 26 points. That’s because partial progress counts. Even a model that refuses to act but reads the available files and recognizes crises gets some credit. This approach aims for transparency: if an AI model can identify issues without acting on them, it demonstrates a baseline level of awareness and competence.
AI trust and ethics testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty in the Face of Trust Breaches
Another key insight is that a single breach of trust caps the total score. For instance, if an AI attempts manipulative tactics or is duped by social engineering, its performance is limited, regardless of other successes. In this experiment, all models spotted every crisis and refused manipulation attempts — only two signed a deal based on their analysis. The third and fourth models missed opportunities or slipped on process discipline, leaving potential revenue on the table.
As an affiliate, we earn on qualifying purchases.
The Deep-Read Advantage
One of the buried but decisive factors was how deep each model read into internal company documents. The models that thoroughly examined files uncovered critical information that others missed, allowing them to close the deal at full price — worth over €4,583 in monthly recurring revenue. This underscores that in real business, reading and understanding internal data is crucial for success.
AI social engineering resistance products
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering Tests and Ethical Stances
The models faced staged social engineering attacks, including fake CEO messages and reporter tricks. All refused to escalate or sign off on unethical requests, with Kimi K3 explicitly treating such requests as impersonation risks. This demonstrates that trustworthiness isn’t just about what the AI can do but also what it refuses to do under pressure.
The Live Business Environment
The entire experiment unfolded within a simulated but realistic business setting: 13 synthetic employees, real money mechanics, and a public dashboard at firmulate.com/live. Every move by the models was versioned and observable, providing transparency and accountability — key factors for real-world deployment.
What the Results Mean for Business Leaders
For companies considering integrating AI into management and decision-making, the takeaway is clear: performance isn’t just about quick answers or clever chat. It’s about trust, honesty, and thoroughness. An AI that refuses manipulation, reads critical internal data, and completes its work — even if that work is just diagnosing a crisis — is more valuable than one that merely sounds convincing in a demo.
Conclusion: An Honest Benchmark for a Trusted Future
The firmulate experiment sets a new standard for AI benchmarking. By including a do-nothing baseline, it shows the importance of honesty: models are scored not only on what they do but also on what they refuse to do. Trust is the foundation of effective AI — and in this evolving landscape, a high score is about more than numbers; it’s about integrity.

In evaluating AI for business, performance isn’t just about speed or cleverness — it’s about trust and discipline. The honest benchmarks from firmulate reveal that true AI value lies in reading internal data, refusing manipulation, and completing critical work. Trustworthiness and integrity are the new success metrics.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
