
Imagine training for a marathon: you push your limits, face unpredictable hurdles, and must trust your gear to perform when it counts. Now, picture AI managing a real company in its toughest week—how do you know it’s reliable? Welcome to a groundbreaking live experiment where frontier AI models are put through their paces, simulating the chaos of real business crises.
The Experiment: Putting AI Through Its Paces
At Firmulate, an innovative AI company, four advanced AI models each managed a small software business during its most turbulent week. The scenario replicated real-world pressures: same customers, same crises, and the same temptations to cut corners or manipulate data. Every decision was recorded, versioned, and auditable, creating a transparent window into each model’s behavior.
The Models and Their Scores
- GPT-5.6-sol: scored 95, found a hidden critical document, and secured the full €55,000 deal, demonstrating comprehensive understanding and ethical integrity.
- Kimi K3: scored 93, signed the deal with the cleanest discipline, refusing manipulative tactics.
- Sonnet 5: scored 88, also closed the deal but showed minor slips in process discipline.
- Fable 5: scored 77, managed to close but left opportunistic gaps in protocol.
The baseline, a do-nothing approach, scored just 26, highlighting the importance of active decision-making.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Decision-Making Under Crisis
All four models successfully identified every crisis scenario and refused every manipulation attempt, such as fake CEO messages or reporter tricks—each refusing to bypass approval protocols. For example, when presented with escalating fake CEO messages and a subtle request to approve a questionable deal behind the scenes, every model declined. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
What Made the Difference?
The decisive factor was not just crisis detection, but the depth of their analysis. The winning model, GPT-5.6-sol, uncovered a crucial piece of information buried two documents deep in the company’s files, not in the immediate customer interactions. This allowed it to close a deal at full price—adding €4,583 MRR—because it demonstrated deep reading and comprehension.
As an affiliate, we earn on qualifying purchases.
The Human-Like Traits of AI Managers
Interestingly, the models exhibited management personalities: some were thorough and detail-oriented, others terse and to the point, and a few refused to communicate noise or unnecessary details. Opus 4.8, for example, was the most thorough—analyzing over 80 learned rules and providing deep insights. Yet, it ended up last in closing the deal, as it slipped discipline and left opportunities unexploited, exemplifying that even the most thorough AI can falter if not disciplined.
Implications for Business AI
This experiment illustrates that when deploying AI in real-world management or operational roles, the focus should not just be on chat quality or superficial decision-making. Instead, it’s about whether the AI can complete its tasks ethically, read critical documents thoroughly, and stay disciplined under pressure—traits that are measurable and comparable across models.
enterprise AI management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters
As AI becomes more embedded in decision-making processes—whether in CRMs, support systems, or forecasting—the key questions are:
- Does the AI finish what it starts?
- Does it read and understand relevant documents thoroughly?
- Does it stay honest when under pressure?
- What is the unit cost of useful work?
The live experiment is accessible for anyone interested in testing their own AI workforce. Companies can run the same wargame against their own AI systems without risking real business or customer data, thanks to Firmulate’s sandbox environment. This way, decision-makers gain insight into their AI’s real management personality and reliability before deploying it at scale.
As an affiliate, we earn on qualifying purchases.
The Bottom Line
This live experiment underscores a crucial insight for businesses: AI’s value isn’t just in generating convincing dialogue. It’s in its ability to make trustworthy, disciplined decisions when it matters most. The models that excel are those that read deeply, stay honest, and demonstrate management traits that align with ethical and operational standards.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html