
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Wellness depends on more than good intentions
In health and wellness, trust is part of the work: people need to know that care, privacy and sound judgment will hold up under pressure. Businesses face a related test as they consider AI. A polished answer in a demo is one thing. What happens when an AI workforce faces a crisis, a tempting shortcut or a chance to close a deal?
Firmulate’s live experiment puts that question to work. It asks models to run the same small software company through its worst week, then makes their decisions available to watch. For organizations considering AI, the next step is to run the test against their own business.
Same crises, different choices
In the final Crucible League, published in July 2026, five participants placed in this order: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The experiment’s standard is uncompromising: “no amount of good work outweighs a breach of trust.”
Each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The surprising result was not that the models missed danger: all spotted every crisis and refused every manipulation attempt. It was that only two signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The detail hidden in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical management question: can an AI system bring relevant information together and then follow through?
The experiment also tested social engineering. Fake CEO messages escalated across three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough analysis is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.
There is a fairness note for readers comparing the standings: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings offer a snapshot of this particular experiment, not a universal promise about how models will perform in every company.
A company you can watch
The live Firmulate company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For wellness organizations, the lesson is concrete: test how AI handles trust, pressure, records and follow-through before giving it a role in sensitive work. A board report should expose the weak points in the playbooks as well as rank the models.

From watching to a pilot
Enterprises can run the same kind of wargame against a read-only export of their own business. The test uses company data and crisis scenarios, then produces a board report with model rankings and the weak points in existing playbooks. Nothing writes back to real systems.
To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
