
In fitness, as in tech, effort alone doesn’t guarantee results. You can put in countless hours of training, yet sometimes miss the essential move that makes all the difference. The same applies to AI models assessing complex situations. The recent experiment by Firmulate sheds light on this paradox: thoroughness and volume don’t always translate into impact.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: Testing AI in a Simulated Business Crisis
Imagine running a small company through its toughest week—same customers, same crises, same temptations. Now, picture four different AI models being tasked to navigate this scenario, decision by decision, with every move recorded and verified. That’s precisely what Firmulate did in its latest live experiment, pitting leading frontier models against one another in a high-stakes business simulation.
The Players and the Stakes
In the crucible league final held in July 2026, four models competed: GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5. The goal was straightforward: can they detect crises, avoid manipulation, and close deals worth €55,000? The baseline was a do-nothing approach, which scored just 26 points, highlighting how easy it is for AI to do nothing and still fall behind.
Key Findings: Attention to Detail Doesn’t Guarantee Impact
- All four models identified every crisis and refused every manipulation attempt—showing a strong sense of honesty and awareness.
- Only two of the four managed to close the deal and sign the contract, despite all having similar diagnoses and pitches.
- The critical weakness was not in the immediate crisis response but in how the models read company files—specifically, information located two document references deep in the company’s files.
- Models that thoroughly read and understood these files were able to win the deal at full price, adding over €4,500 monthly recurring revenue (MRR).
Social Engineering and Integrity
The experiment also tested AI resilience against social engineering: fake CEO messages escalating in urgency and a reporter’s trick asking for a background yes/no response. All five models refused these manipulative requests, with Kimi K3 explicitly citing the risk of impersonation or approval bypass.
The Human-Like Company and Its Challenges
The simulated company comprised 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against €2,300 MRR, with a visible cash countdown. The company operated according to over 680 self-learned rules, with every decision versioned and publicly observable at firmulate.com/live. This setup aimed to mirror a real-world business environment, testing AI’s ability to navigate complex, high-pressure situations.
The Paradox of Diligence and Impact
Among the models, Opus 4.8 stood out for its thoroughness—learning over 80 rules with deep analyses. However, it finished in last place—failing to escalate write attempts into the appropriate departments and leaving the closed deal on the table. This divergence underscores a crucial lesson: meticulousness and volume of learned rules do not necessarily translate into effective impact or decision-making.
Implications for Business and AI Adoption
This experiment highlights a vital consideration for companies integrating AI into decision-making roles: impact depends more on prioritization and focus than sheer volume of information processed. An AI model that reads all documents deeply and understands their relevance is more likely to close deals and maintain integrity than one that merely learns many rules without strategic discipline.
As an affiliate, we earn on qualifying purchases.
What Business Leaders Should Take Away
In the pursuit of smarter AI, the focus must shift from volume to discernment. Effort is important, but without prioritization, even the most diligent models can miss the point. The Firmulate experiment shows that AI models capable of reading deeply and understanding context at the right moments outperform those that focus on exhaustive rule-learning alone.
For enterprises considering AI for critical decision-making, this means deploying systems that prioritize understanding over volume, and integrity over mere compliance. Tests like these—public, transparent, and observable—are invaluable in preparing your business for an AI-powered future where impact counts more than effort.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.