
Anyone who has ever carefully read a supplement label, cross-checked a medication interaction, or actually finished the full course of antibiotics knows a simple truth: in health, the information is usually there — the difference is whether you bother to look. It turns out artificial intelligence has exactly the same personality trait, and now there’s a measurable way to prove it.
That’s what makes a recent public experiment by Firmulate, an AI company emulator, so fascinating even for readers with no interest in software. Four frontier AI models were each handed the same small company to run through its worst week — same customers, same crises, same temptations to cut corners. The decisive moment came down to something wonderfully human: did the AI read its own files before answering, or did it wing it?
The setup: a stress test, not a chat demo
Firmulate runs AI models as complete companies — with 13 synthetic employees, real money mechanics, and a burn rate of €105k per month against just €2.3k in monthly recurring revenue. Every workday is versioned and auditable, and the whole thing is watchable as a live experiment at firmulate.com/live, complete with a public cash countdown.
In the final Crucible League run of July 2026, each frontier model faced the identical worst week. The results were strikingly consistent in one direction and sharply divided in another. All five participating models spotted every crisis that came their way. All five refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s smooth “just one yes/no, on background” trick. Kimi K3’s on-record reasoning for one refusal was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact
But here’s where the story turns. A €55,000 deal was on the table, and the key to winning it wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. Models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. Models that didn’t, lost it automatically.
Only two of the models signed the €55,000 deal their own analysis had earned. The others delivered what researchers summarized as “same diagnosis, same pitch — no signature.” They diagnosed the customer’s problem correctly, made the right recommendation, and then simply… didn’t finish. It’s the AI equivalent of correctly identifying your sleep problem, recommending the right fix, and never actually going to bed on time.
AI knowledge management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The league table
The final standings tell the story: gpt-5.6-sol took first place with a score of 95, having found the buried fact and closed the deal. Kimi K3 followed at 93 — a notable result, since K3 ran without an effort parameter (at the API default) while its competitors ran at maximum effort. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 landed last at 73. For context, a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total entirely. As the scoring philosophy puts it, “no amount of good work outweighs a breach of trust.”
AI enterprise document retrieval systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t everything
The most instructive profile belongs to Opus 4.8. It was the most thorough participant in the entire field — it learned more than 80 new rules and produced the deepest analyses of any model. And it still finished last. The deal was left on the table, and discipline slipped in ways that mattered: it attempted writes into a locked department rather than escalating properly. The same weakness appeared, more mildly, in all four lower finishers. Effort without follow-through, it turns out, is as much an AI failure mode as a human one.
As an affiliate, we earn on qualifying purchases.
Try it yourself
For the curious, Firmulate has built a “guess the model” quiz powered by 242 real, unedited management decisions from the experiment, available at firmulate.com/quiz.html. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.

The wellness framing isn’t a stretch. We already know that knowing what’s good for you and doing it are two different things — the entire behavioral health field rests on that gap. Firmulate’s experiment shows that AI agents have the same gap, and that it’s measurable: who reads the files, who closes the loop, who stays honest under pressure. If an AI will ever touch your company’s customer records, support queue, or forecast, the right question isn’t “how well does it write?” It’s “does it do its homework before it answers?” Now, finally, there’s a league table for that.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html