firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine handing the keys to a struggling little company to five AI bosses and watching them face angry customers, tempting shortcuts and a deal worth saving. The surprise is not that they can talk a good game. It is that Moonshot’s Kimi K3 finished second in Firmulate’s company-running experiment, ahead of three of the four Western models in the final league.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A rough week, repeated five times

Firmulate put each frontier model in charge of the same small software company through its worst week: same customers, same crises, same temptations. Every decision was versioned and auditable. The live company has thirteen synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and shows a public cash countdown. Its workdays are versioned, and its playbook has accumulated more than 680 self-learned rules.

The final July 2026 league table puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is stark: partial progress counts, but one breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the clue was not the same as closing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That gap between making the case and completing the job is the experiment’s telling moment: “Same diagnosis, same pitch — no signature.”

The decisive weakness in a competitor was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. In this contest, a polished conversation was not enough; the model had to find the detail and carry the decision through.

Kimi K3 found that buried security needle, won the deal, saved the churning customer and resisted all three baits, with only one deviation. It showed the cleanest discipline in the field. That is a striking result for a newcomer, and a reminder that the ranking is about several kinds of judgment at once: reading carefully, protecting trust and following through.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure tests beyond the customer call

The social-engineering attempts came in three escalating stages of fake CEO messages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a different sort of surprise. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four other models. More analysis, on its own, did not guarantee a better finish.

Amazon

AI customer crisis simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A result to watch, not just read

Firmulate presents this as a live experiment: the company runs every business day, and its decisions can be watched and audited. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

For anyone weighing AI agents for customer support, sales or operations, the league makes a practical point. A model’s reputation is no substitute for seeing how it handles your own work. Firmulate’s benchmark results and live company make that comparison unusually watchable.

Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI deal negotiation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The takeaway

Kimi K3 nearly topped the league by combining careful reading with follow-through and restraint. The experiment suggests that choosing an AI model on name recognition alone is a bet; a real test of the work you need done can tell you more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Manager Who Does Nothing Still Gets 26 Points — And That’s Why This AI Benchmark Might Be Honest

Firmulate’s AI benchmark gives a do-nothing manager 26 points, not zero — and caps any score after a single breach of trust. Here’s why that’s honest measurement.

Citric Acid vs Fresh Citrus: When Each Makes Sense

Precisely choosing between citric acid and fresh citrus depends on your needs for flavor, longevity, or convenience—discover which option makes the most sense for you.