
In business, the most consequential choice may be the one that looks least dramatic: sign the deal, follow the evidence, and know when a request is a trap. Firmulate turns those moments into a live experiment, asking what happens when AI models have to run a company rather than merely talk about one.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, on repeat
In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, crises and temptations. Their decisions were versioned and auditable. The published order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The experiment’s surprise was not that the models failed to notice danger. Every model spotted every crisis and refused every manipulation attempt. The divide came after diagnosis: only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The clue was already in the company’s files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a plain business point: having the right analysis is not the same as acting on it, and a decisive fact can be easy to overlook even when it is already available.
Trust faced a more direct test. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different kind of caution. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and let discipline slip, making write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. Thoroughness alone did not secure the outcome.
Watch the experiment, then try your own
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. It is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice at firmulate.com.
The next step is to move from watching a company face pressure to testing your own. Enterprises can run the wargame against a read-only export of their business, with crisis scenarios and a board report on model ranking and weak points in their playbooks. Nothing writes back to real systems. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh, a fairness detail to keep in mind when reading the league.

Put your playbooks through a real test
A model that recognizes a crisis still has to make the decision that protects the business and completes the work. Firmulate’s live experiment makes those choices visible; a pilot can put your company’s own plans and pressure points to the test. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
