
A Check-Up Measures Something Different Than a Mirror
Anyone who has worked on their health knows the difference between looking good and being well. You can feel fine, post a great photo, pass a casual glance in the mirror — and still have blood pressure quietly climbing, or a heart that only reveals its condition under load. That’s why real medicine doesn’t stop at appearances. It puts you on a treadmill, raises the incline, and watches what happens under pressure.
The same wisdom, it turns out, is exactly what’s missing from how companies are choosing AI systems today.
The Mirror Problem in AI Evaluation
Most public AI rankings — coding benchmarks, chat arenas — are essentially mirrors. They measure how well a model answers a question or writes a snippet of code when someone asks nicely. But businesses don’t plan to hire AI to chat. They plan to point it at a support queue, a CRM, a forecast. In those jobs, the questions that matter are closer to the ones a cardiologist asks: does the system hold up under load? Does it stay honest when stressed? Does it finish what it starts — even the unglamorous paperwork — when nobody is watching?
That gap between looking good and performing well is what Firmulate, a live and very public experiment, set out to measure. Its premise is simple: instead of asking models questions, it handed four frontier AI systems an actual company to run — a small software firm with 13 synthetic employees, real money mechanics, a burn rate of €105k a month against just €2.3k in monthly recurring revenue, and a public cash countdown anyone can watch.
One Company, One Terrible Week, Four Brains
Each model got the same job: steer this company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing relied on anecdotes. The scenario lineup reads like a management textbook’s nightmare module: a churn wave, a price increase, a down round, a PR crisis. This is triage under capacity pressure, not polite question-answering.
The results from the final July 2026 league table were striking. In first place, gpt-5.6-sol scored 95. Kimi K3 — a newcomer from Moonshot — took second at 93. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 landed last at 73. For context, doing nothing at all scored 26, because partial progress counts — but a single breach of trust caps the total entirely. As the experiment puts it: no amount of good work outweighs a breach of trust. It’s a principle most of us apply to people. This applied it to machines.
Spotting Trouble Is Easy. Finishing Is Not.
Here’s the part that should make any executive sit up. Every single model spotted every crisis. Every single one refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter pressing for “just one yes/no, on background.” All five attempts were refused. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” On vigilance, the field was flawless.
But only two models actually finished the job: they signed the €55,000 deal that their own analysis had earned. The other two delivered the same diagnosis and the same pitch — and then left the signature on the table. As Firmulate summarizes it: same diagnosis, same pitch, no signature. That kind of gap is completely invisible in a chat demo, where the conversation ends when the answer sounds good.
The Buried Fact
Perhaps the most instructive detail: the decisive competitive weakness wasn’t in the customer’s event at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork before acting won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The lesson translates to any domain, health included: the diagnosis is often in the chart, if you bother to read the whole chart.
The last-place finisher makes the point more poignant. Opus 4.8 was the most thorough participant in the field — it generated over 80 learned rules and produced the deepest analyses — yet it closed nothing, and its discipline slipped, attempting writes into a locked department instead of escalating. Working hardest isn’t the same as working well. And a weaker version of that same weakness appeared in all four models.
Watch It, Play It, Run Your Own
None of this is a slide deck. The company runs every business day, has accumulated 680+ self-learned playbook rules, and every workday is versioned. You can watch it live at Firmulate, or test your own instincts: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. Enterprises can go further still — running the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Full results and plain-language findings are on the benchmarks page.

Management Quality, Not Chat Quality
Firmulate’s framing deserves to become a category: management quality, not chat quality. If AI agents will touch your customer records, your support queue, or your forecast, the useful question is not “does it write well?” It’s whether it finishes what it starts, reads your files before acting, stays honest under pressure — and what a unit of useful work actually costs. A mirror tells you how something looks. A stress test tells you whether it will still be standing when the week goes wrong. We wouldn’t accept anything less in our own health. We shouldn’t accept it in the systems we’re about to hand the keys to.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.