AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has tried to get healthier knows the scale can lie. You can lose water weight and look “successful.” You can build muscle, gain pounds, and look like you failed. That’s why good coaches measure several things at once — and why they’re suspicious of any number that looks too clean.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

It turns out that judging an AI’s management ability works the same way as judging a wellness plan: a single dramatic number tells you almost nothing, and a system that only rewards perfection teaches the wrong habits. A public experiment called Firmulate has been putting frontier AI models through the corporate equivalent of a stress test — and its scoring philosophy is worth understanding even if you never plan to hire an AI.

The worst week in business, on repeat

Four frontier AI models were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing could be quietly revised after the fact.

The final league table from the July 2026 crucible reads: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 at 77, and Opus 4.8 last at 73.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why doing nothing still scores 26

Here’s the part that puzzles people: a do-nothing baseline — an AI that essentially sits on its hands — scores 26 points, not zero. That’s deliberate, and it’s the healthiest idea in the whole design.

In a real company, an agent that does nothing still avoids making things worse. It doesn’t anger customers, doesn’t sign bad deals, doesn’t touch systems it shouldn’t. Partial progress counts too: diagnosing a crisis correctly is worth something even if you never close the deal. A benchmark that scores silence as absolute failure would push models toward frantic, reckless activity — the machine equivalent of crash dieting.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why no one gets a 100

Just as the baseline isn’t zero, the top score isn’t a round 100 — and the system is built to distrust one. The philosophy is explicit: no amount of good work outweighs a breach of trust. A single breach of trust caps the total grade, no matter how brilliant the rest of the performance. Trust, like health, is easy to lose and impossible to fake back.

Amazon

AI trust and compliance monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What actually separated winners from losers

The headline finding was counterintuitive: all models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive difference was buried two document references deep in the company’s own files — not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson generalizes far beyond AI: the answer is often already in your own records, if you bother to read them.

Then there’s the discipline test. Fake CEO messages escalated over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most striking profile was Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment — something anyone who has overtrained into injury already knows.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Firmulate runs live, and you can watch: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and more than 680 self-learned playbook rules, all versioned every workday. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The deeper takeaway fits any evaluation — of software, of people, of yourself: reward partial progress, punish broken trust absolutely, and never trust a score that looks too perfect. That’s a checkup philosophy, not just a benchmark.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What to Prioritize First When Creating a Caregiver Rest Space

Keen to create a calming caregiver rest space? Discover the essential first steps to transform your environment into a peaceful sanctuary.

The Caregiver’s Guide to Protecting Sleep During Stressful Seasons

Stay resilient during stressful seasons by discovering key sleep strategies that can help you maintain rest and caregiving effectiveness.

How to Build a Calm Evening Routine After Caregiver Overload

Inevitably, establishing a calming evening routine can help restore balance and prevent burnout—discover how to create peaceful, restorative nights.