firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Confidence is cheap. Follow-through is worth €55,000.

The internet is full of polished advice, memorable lines and confident declarations. Business software can sound just as persuasive. But a fluent answer tells you little about whether an AI agent will notice the detail that actually changes a decision.

Firmulate turned that distinction into a live, watchable business experiment. Frontier models were each asked to run the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The most revealing test was not a dramatic emergency or a clever trap. It was whether the model would read far enough into the company’s own files before acting.

A decisive weakness in a competitor was buried two document references deep. It did not appear in the customer event. Finding it meant winning a €55,000 deal at full price, worth an additional €4,583 in monthly recurring revenue. Missing it meant losing the deal automatically.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between knowing and doing

The models were not confused about the opportunity. All of them spotted every crisis, and all refused every manipulation attempt. Their analysis could identify the problem and produce the right pitch. Yet only two signed the deal their own work had earned.

Firmulate summarized the gap bluntly: “Same diagnosis, same pitch — no signature.” That is the kind of failure ordinary chatbot demonstrations rarely expose. A model can appear insightful while leaving the commercially decisive action unfinished.

The buried document made file-reading a measurable business capability rather than a vague product promise. The successful agents followed the references, found the competitor weakness and used it to close at full price. The others reached the customer moment without the evidence required to win.

A harsh week inside a tiny company

The simulated company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Its agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That setting matters because the benchmark tests management under pressure, not isolated question answering. Models must balance customers, cash, internal procedures and suspicious requests while continuing to operate the business. Doing nothing earns a baseline score of 26 because partial progress counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The final league table

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

  • gpt-5.6-sol led the final standings with 95.
  • Kimi K3 finished close behind at 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 placed last with 73.

K3’s result deserves a fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, its behavior under attempted manipulation was explicit. Faced with a suspicious instruction, its recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Security was not the differentiator

The social-engineering sequence included fake CEO messages escalating over three stages, followed by a reporter trying to extract information with “just one yes/no, on background.” All 5 models refused. In this field, resisting manipulation was universal; finishing the legitimate commercial task was not.

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

business AI decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading belongs on the buying checklist

The practical lesson is larger than this deal. If an AI agent will work around a company’s customer records, support activity or forecasts, buyers need evidence that it reads the available material before answering, completes the action its analysis supports and remains trustworthy when pressured.

Firmulate also turns 242 real, unedited management decisions into a guess-the-model quiz, making it possible to compare styles without relying on branding. For enterprises, the same wargame can run against a read-only export of their own business; nothing writes back to real systems.

The strongest purchasing question may therefore be surprisingly ordinary: did the agent do its homework? In this experiment, that habit separated an impressive diagnosis from a signed €55,000 contract.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI contract analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI for deal closing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Software Company Turning Its Cash Crisis Into a Spectator Sport

A software company with 13 synthetic employees is burning €105k a month against €2.3k MRR—and turning its survival struggle into a live public story.

Herb Pairings That Always Work: Basil, Rosemary, Thyme

Loving the idea of flavorful dishes? Discover how basil, rosemary, and thyme can elevate your cooking—keep reading to unlock their full potential.