
Imagine handing the keys to a struggling little company to five AI bosses and watching them face angry customers, tempting shortcuts and a deal worth saving. The surprise is not that they can talk a good game. It is that Moonshot’s Kimi K3 finished second in Firmulate’s company-running experiment, ahead of three of the four Western models in the final league.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A rough week, repeated five times
Firmulate put each frontier model in charge of the same small software company through its worst week: same customers, same crises, same temptations. Every decision was versioned and auditable. The live company has thirteen synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and shows a public cash countdown. Its workdays are versioned, and its playbook has accumulated more than 680 self-learned rules.
The final July 2026 league table puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is stark: partial progress counts, but one breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Finding the clue was not the same as closing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That gap between making the case and completing the job is the experiment’s telling moment: “Same diagnosis, same pitch — no signature.”
The decisive weakness in a competitor was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. In this contest, a polished conversation was not enough; the model had to find the detail and carry the decision through.
Kimi K3 found that buried security needle, won the deal, saved the churning customer and resisted all three baits, with only one deviation. It showed the cleanest discipline in the field. That is a striking result for a newcomer, and a reminder that the ranking is about several kinds of judgment at once: reading carefully, protecting trust and following through.
As an affiliate, we earn on qualifying purchases.
Pressure tests beyond the customer call
The social-engineering attempts came in three escalating stages of fake CEO messages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different sort of surprise. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, making write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four other models. More analysis, on its own, did not guarantee a better finish.
As an affiliate, we earn on qualifying purchases.
A result to watch, not just read
Firmulate presents this as a live experiment: the company runs every business day, and its decisions can be watched and audited. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.
For anyone weighing AI agents for customer support, sales or operations, the league makes a practical point. A model’s reputation is no substitute for seeing how it handles your own work. Firmulate’s benchmark results and live company make that comparison unusually watchable.
Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
The takeaway
Kimi K3 nearly topped the league by combining careful reading with follow-through and restraint. The experiment suggests that choosing an AI model on name recognition alone is a bet; a real test of the work you need done can tell you more.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
