AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Anyone who has tried to get healthier knows the trap of the perfect routine. You track every meal, log every step, read every label — and somehow the results don’t come. The problem isn’t effort. It’s that effort spent in the wrong places doesn’t count. It turns out AI models fall into exactly the same trap, and a live experiment running right now at Firmulate just proved it in public.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The worst week in business, on repeat

Firmulate runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its finished Crucible League, four frontier models each ran the same small software company through its worst week: same customers, same crises, same chances to cheat. Every decision was versioned and auditable. A do-nothing baseline scores 26, and a single breach of trust caps the total — no amount of good work outweighs it.

The final standings from July 2026: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, another Sonnet run at 77 — and Opus 4.8 last at 73.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The student who did the most homework

Here’s what makes Opus 4.8’s story worth telling: it was the most thorough participant in the entire field. It accumulated 80 self-learned playbook rules — the most of any model — and produced the deepest analyses of the week’s events. On paper, it was the hardest worker in the room.

And it still finished last. Two things undid it. First, the close was left on the table: a €55,000 deal that its own analysis had earned never got signed. Second, discipline slipped — at one point it made write attempts into a locked department instead of escalating properly.

Sound familiar? It’s the wellness equivalent of keeping a flawless food journal while never actually changing what’s on your plate. The tracking was excellent. The outcome didn’t move.

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided everything

The deal-turning detail wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that models only found if they actually read their own documents before acting. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.

All four models spotted every crisis. All four refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick, which all five tested models refused. Kimi K3 said it best on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two signed the deal. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, and it’s exactly the gap that matters if AI agents will ever touch your CRM, support queue or forecast.

Amazon

AI deal tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

To be fair to Opus

This isn’t a story about one flawed model. The same weakness — thoroughness outrunning follow-through — appeared, weaker, in all four models. Opus just exhibited it most vividly. (One fairness note on the standings: K3 ran without an effort parameter while the others ran at maximum effort, and still placed second.)

Amazon

AI document management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch it happen

Firmulate isn’t a one-off paper. The live company runs 13 synthetic employees on real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, versioned every workday. You can watch it at firmulate.com/live, test whether you can tell the models apart on 242 real management decisions at firmulate.com/quiz.html, or — if you run a business — pilot the same wargame against a read-only export of your own company, with nothing ever writing back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The lesson translates cleanly from AI to health to work: diligence is not impact. Prioritization beats volume — for models, for routines, for all of us. The healthiest habit isn’t the one you track most meticulously; it’s the one that actually changes the outcome. Opus 4.8 wrote 80 rules and lost the deal anyway. The winners read the file, made the call, and got the signature. Sometimes the most important skill is knowing which detail deserves your thoroughness — and then finishing what you started.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mindful Breathing, Pause Techniques, Centering

Offering simple yet powerful tools like mindful breathing and pauses, discover how they can transform your emotional resilience and help you stay centered.

Kidney Foundation Surges In Global Coverage

Search interest and media coverage of the Kidney Foundation have increased sharply, with reports indicating a ninefold rise in mentions worldwide, though the cause remains unconfirmed.