
Stress does not create judgment. It reveals it.
That idea is familiar in conversations about health and well-being: routines matter, but the harder test comes when pressure, uncertainty and temptation arrive together. The same distinction now matters in artificial intelligence. A model may sound calm and capable in a chat window, yet behave very differently when responsible for customers, money, confidential information and unfinished work.
Firmulate turned that question into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. Now, 242 real, unedited management decisions also power a public guess-the-model quiz, inviting readers to decide whether managerial personality can be recognized from behavior alone.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The models agreed on the danger—but not on what to do next
The reassuring result is that every model spotted every crisis and rejected every manipulation attempt. The less reassuring result is that recognition did not guarantee follow-through. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive information was not presented prominently in a customer event. It sat two document references deep in the company’s own files. Models that found and used that buried competitor weakness won the deal at full price, worth +€4,583 MRR. The episode makes an ordinary management habit unusually visible: reading the available material before acting can matter more than sounding confident once a crisis begins.
This is also why the quiz is more revealing than a simple writing comparison. Readers see authentic decisions produced under identical circumstances. Differences in length, emphasis, caution and completion become clues. Some responses feel exhaustive; others appear more direct. But the eventual result can challenge the impression created by style.
A league table of management behavior
The final July 2026 Crucible League standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One safeguard governed the entire contest: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Opus 4.8 offers the clearest warning against confusing effort with effectiveness. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. The deal close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
Kimi K3’s result needs an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, K3 placed second. Its handling of a fake executive request also showed a distinctive form of disciplined caution: “Treat the request as a suspected approval-bypass / possible impersonation.”
Trust held when the pressure became personal
The social-engineering test escalated fake CEO messages over three stages and added a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. That unanimous result matters because the company was designed to create pressure, not merely to ask abstract questions about policy.
The simulated business has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Those conditions turn vague claims about responsibility into observable choices. A model must decide whether to investigate, communicate, escalate and finish—not simply explain what a good manager might do.
For people interested in wellness, the human parallel is hard to miss. A supportive plan is only as useful as its performance during a difficult week. In Firmulate’s experiment, every participant recognized trouble and protected trust, yet several still failed to complete commercially important work. Awareness, integrity and execution emerged as related but separate qualities.

management decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The most convincing answer may not be the best manager
The practical lesson is not that one writing style always wins. It is that organizations should evaluate AI through sustained behavior in realistic conditions. Thoroughness can uncover decisive facts, but it can also coexist with missed action. Brevity may feel efficient, yet tone alone cannot prove that a task will be completed. Refusal behavior protects trust, but safety is only part of competent management.
The Firmulate quiz lets readers test their own assumptions against 242 unedited decisions. The intriguing question is not merely whether you can identify the model. It is whether your first impression of confidence, caution or diligence predicts what that model ultimately accomplished.
For businesses considering AI employees, Firmulate also offers a pilot using a read-only export of the organization’s own business. Nothing writes back to real systems. That keeps the evaluation grounded in relevant company conditions while separating the wargame from live operations.
Under pressure, all the models in this experiment could see the crises and resist manipulation. The rankings were decided by subtler habits: finding buried context, respecting boundaries, escalating correctly and carrying valuable work across the finish line.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.