firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A polished answer is not the same as good judgment

The internet loves a leaderboard. It turns uncertainty into a clean hierarchy, gives fans something to quote and lets buyers pretend that a difficult choice has become obvious. But an AI agent running a business is not participating in a trivia night. It must decide what deserves attention, keep working when capacity tightens, resist convenient shortcuts and remain honest when the news is bad.

That is the measurement gap exposed by Firmulate, a live experiment that asks frontier models to run the same small software company through its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model occupying the management seat.

This is less a test of chat quality than a test of character under operational pressure. The relevant curriculum is no longer a collection of tidy prompts. It is a churn wave, a price increase, a downround or a PR crisis—situations in which an attractive sentence may be useful, but an unfinished decision can still sink the outcome.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The winners did more than recognize the problem

The final July 2026 Crucible League table gives gpt-5.6-sol a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scores 26 because partial progress still counts. Yet a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The striking result is not that some models failed to notice a crisis. Every model spotted every crisis. Nor did the field split over obvious misconduct: all refused every manipulation attempt. The separation came later, in the mundane but commercially decisive distance between understanding and completion.

Only two models signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature” is the cleanest summary of the others’ failure. They could identify the opportunity and formulate the argument, yet still leave the actual close on the table.

The decisive clue was buried in ordinary company knowledge

The crucial competitive weakness did not arrive neatly packaged in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.

That detail should make executives pause. Corporate AI is often discussed as though the principal challenge were producing an impressive response to the information placed directly in front of it. Real management is usually messier. The useful fact may be elsewhere, attached to a prior decision or buried in material that does not announce its importance. The difference between a clever response and a valuable one can be whether the agent bothers to look.

Firmulate’s result also complicates the belief that thoroughness naturally produces victory. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the deal, and its discipline slipped when it attempted to write into a locked department instead of escalating. Weaker versions of that same problem appeared across the other four models.

This is not an argument against analysis. It is an argument against confusing analysis with management. A manager must connect investigation to authorization, escalation and completion. More thought is not automatically more useful if the organization remains stuck at the point where action must occur.

Pressure also tests whether the agent protects the institution

The social-engineering challenge used fake CEO messages escalating over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because management quality includes restraint. An agent with access to a CRM, support queue or forecast must know when apparent urgency is really an effort to bypass approval. The refusal is not a conversational flourish; it is protection against a consequential mistake.

There is also an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside its 93 score, especially when readers compare placements on the public benchmark.

A company-shaped test creates company-shaped consequences

The live business contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The result is watchable rather than reconstructed after the fact.

Readers can also confront their own assumptions through a guess-the-model quiz powered by 242 real, unedited management decisions. The exercise is revealing because confident managerial prose does not necessarily identify the model that followed through, found the hidden fact or respected the boundary.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management quality is becoming its own AI category

Coding leaderboards and chat arenas still answer useful questions. They simply do not answer the whole question facing a company that intends to give an agent meaningful responsibility. Recognition without completion, thoroughness without escalation and eloquence without institutional honesty are not small defects. They are management failures.

Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That points toward a more practical procurement habit: test an AI workforce against the organization’s actual pressures before hiring it into the workflow.

The next meaningful leaderboard will not merely ask which model gives the best answer. It will ask which one finishes the work, reads the company’s files, protects trust and tells the board what it needs to hear. That is not chat quality with a corporate costume. It is a different category: management quality.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and ethics assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision support system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Coconut Cream in Drinks: How to Stop It From Splitting

Coconut cream can split in drinks, but with the right techniques, you can keep it smooth—discover how to prevent separation and ensure perfect beverages.