AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

We Take Our Bodies for Annual Checkups. What About the AI We’re Hiring?

Anyone who reads this site knows the routine: you don’t wait for a heart attack to check your blood pressure. You stress-test your health before a crisis finds you. A growing experiment called Firmulate is doing exactly that — not for humans, but for the AI models that businesses are rushing to put in charge of customers, deals, and money.

And the latest result from its “Crucible” league reads like a perfect checkup with one surprise: Moonshot’s Kimi K3, a newcomer, placed second with a score of 93 — ahead of three of the four Western frontier models it competed against.

Amazon

AI stress testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week, Different AI

Firmulate’s premise is simple and a little unnerving. Instead of asking an AI how it would handle a business crisis, it makes the model actually do it. Each frontier model ran the same small software company through its worst week — the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly airbrushed afterward.

The final July 2026 league table tells the story:

  • gpt-5.6-sol — 95 (first place, described as “the complete performance”)
  • Kimi K3 — 93 (closed the deal too, with the cleanest discipline of the field)
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.” That’s a standard most of us would want from our own doctors, and apparently a reasonable one for AI managers too.

Amazon

business AI decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Stress Test Revealed

The findings have a familiar shape for anyone who’s had a thorough checkup: most things were fine, but the weak spots were revealing.

Every model spotted every crisis, and every one refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s trick question framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was strikingly prudent: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two of five models finished the job. The decisive competitive weakness was buried two document references deep in the company’s own files — not in the customer conversation. The models that actually read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered the same diagnosis and the same pitch, but never got a signature. “Same diagnosis, same pitch — no signature,” as the finding goes.

K3 found the buried fact, closed the deal, saved a churning customer, and deviated from protocol only once — the cleanest discipline in the field.

Amazon

AI model performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thorough Isn’t the Same as Healthy

The most human finding in the whole experiment: Opus 4.8 was the most thorough participant, with the deepest analyses and the most learned rules added to its playbook. It also finished last. The close was left on the table, and discipline slipped — the model attempted writes into a locked department instead of escalating. Firmulate noted the same weakness, in weaker form, in all four other models.

It’s the equivalent of a patient with perfect lab results who never actually changes their habits. Effort and follow-through are different muscles — and only a test under real conditions shows which one a model has.

Amazon

AI cybersecurity and trust audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Real, and You Can Watch

This isn’t a thought experiment on a slide. The simulated company has 13 synthetic employees, real money mechanics — a burn rate of €105k per month against €2.3k in monthly recurring revenue — and a public cash countdown. It has accumulated over 680 self-learned playbook rules, and every workday is versioned. You can watch it unfold at firmulate.com, and dig into the full results and plain-language findings on the benchmarks page.

A note on fairness, disclosed by the researchers: Kimi K3 ran without an effort parameter (the API default) while the other models ran at the “xhigh” effort setting — meaning the newcomer’s near-top score came with arguably less wind in its sails.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Test Before You Trust

The uncomfortable conclusion for anyone choosing an AI system in 2026: the league is genuinely open. A newcomer beat three of four Western frontier models at running a company under pressure. Brand reputation no longer predicts performance on the things that matter — finishing what you start, reading the file before the meeting, staying honest when no one is checking.

If AI agents will touch your CRM, your support queue, or your forecast, the question is not “does it write well.” It’s whether it completes useful work under pressure. Chat demos can’t show you that. Only your own test can — and picking a model without one is now a bet, not a decision.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Women Health Surges In Global Coverage

Search interest and media coverage on women’s health have increased sharply worldwide, driven by rising awareness and policy focus, though the exact trigger remains unconfirmed.

Ani Pharmaceuticals Surges In Global Coverage

Ani Pharmaceuticals experiences a significant surge in international media attention, with 24 mentions in recent coverage, indicating growing global interest.

Seniorsplus Lewiston Surges In Global Coverage

Coverage of Seniorsplus Lewiston has surged globally, with 18 mentions in recent media analysis, signaling increased interest in the organization’s activities.