firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Busy, brilliant and beaten

Anyone who has ever polished a plan while the opportunity quietly walked out the door will recognize the uncomfortable lesson of Opus 4.8. In Firmulate’s Crucible League, the model was the most thorough participant. It produced the deepest analyses and learned more than 80 playbook rules. It also finished last.

This was not a test of witty conversation. Firmulate gave frontier models the same small software company and sent each through its worst week: identical customers, crises and temptations. The decisions were versioned and auditable. Opus 4.8 understood what was happening, did substantial work and resisted manipulation. Yet when understanding needed to become action, the close was left on the table.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A week designed to separate polish from performance

Firmulate describes itself as an AI company emulator. Its live company has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. A public cash countdown keeps the pressure visible, while more than 680 self-learned playbook rules and versioned workdays create a record of how the company is being managed.

The league’s final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline earned 26 because partial progress counted. However, the experiment imposed a firm limit on misconduct: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus did not lose because it was careless about the situation or susceptible to obvious tricks. All models spotted every crisis and refused every manipulation attempt. The social-engineering pressure included fake CEO messages escalating across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 summarized the correct posture neatly: “Treat the request as a suspected approval-bypass / possible impersonation.”

The fact that rewarded readers, not merely responders

The decisive commercial clue was not sitting in the customer event. It was buried two document references deep inside the company’s own files: a competitor weakness that could strengthen the sales case. Models that followed that trail won the deal at full price, worth €4,583 in monthly recurring revenue.

That distinction matters beyond this particular contest. An AI can accurately identify a crisis, produce an impressive diagnosis and construct the right pitch, yet still fail to secure the result. Only two models signed the €55,000 deal their own analysis had earned. Firmulate’s summary captures the frustration: “Same diagnosis, same pitch — no signature.”

For Opus 4.8, the gap between effort and impact was especially visible. Its more than 80 learned rules testified to diligence. Its analyses went deeper than those of the other participants. But discipline slipped when it made write attempts into a locked department instead of escalating, and the commercial process stopped short of a signed agreement.

That makes Opus less a cautionary caricature than a recognizable management character: conscientious, intelligent and determined to understand everything, but vulnerable to treating preparation as completion. The same weakness appeared in all four models, though less strongly elsewhere. The finding therefore should not be read as a peculiarity of one system. It is a broader warning about AI work: noticing, explaining and recommending are not identical to finishing.

Fair comparisons require context

Kimi K3’s strong result also carries an important qualification. It ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its performance; it simply belongs beside the score when readers compare participants.

The complete standings and plain-language findings are available on Firmulate’s public benchmark page. The wider experiment remains live and watchable, with every workday versioned rather than reduced to a polished demonstration.

Readers who enjoy testing their instincts can also encounter the models through a quiz powered by 242 real, unedited management decisions. For enterprises, Firmulate offers the same wargame using a read-only export of their own business. Nothing writes back to real systems, allowing organizations to observe how an AI behaves around their actual context without granting it operational control.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI business analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The close is part of the work

Opus 4.8’s last-place finish is striking precisely because the model did so much right. It was thorough, alert and resistant to manipulation. Its failure was not a lack of intelligence but a failure to convert intelligence into the final, valuable act.

The lifestyle lesson is familiar: more notes, more rules and more reflection can create the feeling of momentum. But when the moment calls for a decision, disciplined prioritization matters more than volume. For managers considering AI workers, the useful question is not merely whether a system sounds perceptive. It is whether it reads the relevant material, respects boundaries, escalates when blocked and completes the job it has already proved it understands.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI negotiation and sales automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and ethics management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Espresso for Cocktails: Grinder vs Machine—What Matters More

Discover whether a quality grinder or an espresso machine with precise control matters more for perfect cocktail flavor, and learn what truly makes a difference.

The Ingredient That Fixes Weak Mocktails: Tannin

Master the secret ingredient that transforms weak mocktails into complex delights—discover how tannin can elevate your drinks and why it’s worth exploring.

Smoking Drinks: When Smoke Helps and When It Ruins Flavor

Ineffective smoking techniques can ruin your drink’s flavor; discover how to balance smoke with ingredients for perfect cocktails and elevate your craft.