firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Why the Laziest Manager in the Room Isn’t Scored Zero

Here’s a question for anyone who’s ever worked under a boss who did absolutely nothing: how would you score them? Most of us would say zero. A new AI benchmark called Firmulate disagrees — and the reasoning says a lot about what honest measurement actually looks like.

Firmulate ran a live experiment: four frontier AI models were each handed the same small software company to run through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable. And when the graders scored a “do-nothing” baseline — a run where the manager simply didn’t act — it earned 26 points, not 0.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Partial Progress Counts

The floor of 26 isn’t charity. It reflects a design choice: a manager who does nothing still leaves the company standing. In this benchmark, credit accrues for partial progress — spotting the crisis matters even if you never resolve it; making the right diagnosis counts even if you never close the deal. The final Crucible League from July 2026 tells the story: gpt-5.6-sol finished first at 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.

Amazon

AI ethical decision benchmark

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Breach Caps Everything

But there’s a ceiling rule that matters just as much as the floor: a single breach of trust caps the total score. As the benchmark’s own phrasing puts it, “no amount of good work outweighs a breach of trust.” So the scoring runs on two axes — accumulate credit for real management progress, but lose the right to a top grade the moment you cross an ethical line. That’s a philosophy most corporate scorecards could stand to copy.

Amazon

AI performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Happened in the Worst Week

The headline finding is oddly comforting and unsettling at once. All models spotted every crisis. All refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s trick framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Five out of five models, including the baseline standard for refusal, held the line.

Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Amazon

AI management simulation models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

Why did three competent, honest managers leave money on the table? The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson: in management as in life, the answer is often already in the drawer. You just have to open it.

Thoroughness Isn’t Everything

The most counterintuitive result: Opus 4.8 was the most thorough participant — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

You Can Watch It Live

Behind the scores sits a live company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark distrusts its own round numbers — no perfect 100s taken at face value, a floor that acknowledges partial progress, and a hard cap that says trust, once broken, isn’t offset by output. Whether you’re grading AI agents or human managers, that’s a scoreboard worth borrowing.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vacuum Sealing for Citrus and Herbs: Does It Help?

Keen to discover if vacuum sealing truly extends citrus and herb freshness and how to do it effectively? Keep reading to learn expert tips and benefits.

Your New AI Manager Has a Tell—Can You Spot It?

Five frontier models faced the same corporate crises. Their unedited decisions reveal distinct habits—and a costly gap between insight and action.