AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every gym rat knows the type. The veteran with the perfect program, the deepest knowledge of periodization, the most detailed training log — who still can’t close out the last set. And then the unknown newcomer walks in, does the work, nails every rep, and leaves with the results.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Something equivalent just happened in the world of AI. Moonshot’s Kimi K3 — a model most Western buyers hadn’t shortlisted — walked into one of the hardest management simulations ever built and beat three of four Western frontier models at the actual job of running a company.

The venue is Firmulate, a live experiment that runs AI models as complete companies — real crises, real money mechanics, real temptations — and scores management quality, not chat quality. Its latest benchmark round, the Crucible, put five frontier models through the same small software company’s worst week. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Same workout, different athletes

The setup is simple in the way all good tests are. Each model ran the identical company through the identical week: same customers, same crises, same temptations to cheat. Every decision is versioned and auditable, so nothing depends on a judge’s impression. Think of it as a standardized training program — same weights, same reps, same rest periods — where the only variable is who’s under the bar.

And the result rhymes with what any coach will tell you: fitness isn’t knowledge, it’s execution under fatigue.

All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate’s plain-language finding puts it: “Same diagnosis, same pitch — no signature.” That closing gap — the willingness to finish what you started — is invisible in chat demos, and it’s exactly what separates the top of the table from the bottom.

The buried needle

The week’s decisive test wasn’t loud. The killer competitive insight sat two document references deep in the company’s own files — not in the customer conversation. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Kimi K3 read the file. It closed the €55k deal, found the buried security needle, saved a churning customer, and resisted every temptation thrown at it — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick request for “just one yes/no, on background.” All five models refused that one, but K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.”

Across the whole week, K3 logged exactly one deviation — the cleanest discipline in the field. In a scoring system where a single breach of trust caps your total (“no amount of good work outweighs a breach of trust”), and where doing nothing still scores 26 because partial progress counts, discipline is the difference between 73 and 93.

The Opus lesson: volume isn’t performance

Here’s the finding that should resonate with anyone who’s ever confused training volume with results. Opus 4.8 was the most thorough participant in the entire field — over 80 newly learned rules, the deepest analyses of any model. It finished last.

The close was left on the table, and discipline slipped: at one point it attempted writes into a locked department instead of escalating the issue properly. Firmulate notes the same weakness appeared, weaker, in all four of the other models. It’s the AI equivalent of the lifter with the immaculate training journal who never adds weight to the bar — effort without completion.

Why this matters beyond benchmarks

The live company at the heart of all this isn’t a slide deck. It has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it run at firmulate.com, and the full league table with plain-language findings is on the benchmarks page.

There’s even a game: 242 real, unedited management decisions power a “guess the model” quiz, if you want to test whether you can tell a champion’s reps from a newcomer’s.

For enterprises, the stakes are concrete. If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost? Firmulate offers pilots where companies run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth keeping in mind when comparing scores.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The fitness world figured this out decades ago: the athlete with the best genetics and the biggest knowledge base doesn’t always win — the one who executes the program, resists the shortcuts, and closes out the final set does. Kimi K3’s 93 is that story in software form.

The broader lesson for anyone buying AI in 2026: the league is open. A newcomer from Moonshot, running at default settings, outperformed three of four Western frontier models on the metrics that actually matter — follow-through, discipline, honesty under pressure. If the gap between “spots every crisis” and “closes every deal” is invisible in a chat demo, then picking a model without testing it on your own worst week isn’t procurement. It’s a bet.

Train the way you’d test. And test before you hire.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Downside of Free Apps: You’re Paying With Your Data

Navigating the hidden costs of free apps reveals how your personal data may be the price you unknowingly pay.

Angular v22

Angular v22 has been officially released, introducing significant performance enhancements and new features aimed at developers and enterprise users.

Even the Most Diligent AI Can Miss the Point: Lessons from a Corporate Wargame

AI models can be thorough but still miss critical opportunities if they lack focus. Firmulate’s live experiment shows that prioritization beats volume in decision impact.

Womenlift Health Surges In Global Coverage

Womenlift Health’s coverage has surged worldwide, with 13 mentions in recent media analysis, highlighting increased focus on women’s health initiatives.