
Anyone who has tried to get healthier knows the trap of the perfect routine. You track every meal, log every step, read every label — and somehow the results don’t come. The problem isn’t effort. It’s that effort spent in the wrong places doesn’t count. It turns out AI models fall into exactly the same trap, and a live experiment running right now at Firmulate just proved it in public.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The worst week in business, on repeat
Firmulate runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its finished Crucible League, four frontier models each ran the same small software company through its worst week: same customers, same crises, same chances to cheat. Every decision was versioned and auditable. A do-nothing baseline scores 26, and a single breach of trust caps the total — no amount of good work outweighs it.
The final standings from July 2026: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, another Sonnet run at 77 — and Opus 4.8 last at 73.
As an affiliate, we earn on qualifying purchases.
The student who did the most homework
Here’s what makes Opus 4.8’s story worth telling: it was the most thorough participant in the entire field. It accumulated 80 self-learned playbook rules — the most of any model — and produced the deepest analyses of the week’s events. On paper, it was the hardest worker in the room.
And it still finished last. Two things undid it. First, the close was left on the table: a €55,000 deal that its own analysis had earned never got signed. Second, discipline slipped — at one point it made write attempts into a locked department instead of escalating properly.
Sound familiar? It’s the wellness equivalent of keeping a flawless food journal while never actually changing what’s on your plate. The tracking was excellent. The outcome didn’t move.
business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided everything
The deal-turning detail wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that models only found if they actually read their own documents before acting. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
All four models spotted every crisis. All four refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick, which all five tested models refused. Kimi K3 said it best on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two signed the deal. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, and it’s exactly the gap that matters if AI agents will ever touch your CRM, support queue or forecast.
As an affiliate, we earn on qualifying purchases.
To be fair to Opus
This isn’t a story about one flawed model. The same weakness — thoroughness outrunning follow-through — appeared, weaker, in all four models. Opus just exhibited it most vividly. (One fairness note on the standings: K3 ran without an effort parameter while the others ran at maximum effort, and still placed second.)
As an affiliate, we earn on qualifying purchases.
You can watch it happen
Firmulate isn’t a one-off paper. The live company runs 13 synthetic employees on real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, versioned every workday. You can watch it at firmulate.com/live, test whether you can tell the models apart on 242 real management decisions at firmulate.com/quiz.html, or — if you run a business — pilot the same wargame against a read-only export of your own company, with nothing ever writing back to real systems.

The lesson translates cleanly from AI to health to work: diligence is not impact. Prioritization beats volume — for models, for routines, for all of us. The healthiest habit isn’t the one you track most meticulously; it’s the one that actually changes the outcome. Opus 4.8 wrote 80 rules and lost the deal anyway. The winners read the file, made the call, and got the signature. Sometimes the most important skill is knowing which detail deserves your thoroughness — and then finishing what you started.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.