firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management style is becoming a spectator sport

Personality quizzes usually ask what kind of leader you are. Firmulate turns that idea around: can you identify an AI model from the way it handles a customer, challenges an order or fails to finish an important job?

The premise is unusually concrete. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. Now, 242 real, unedited management decisions power a public guess-the-model quiz. What looks like a diversion quickly becomes a revealing test of judgment.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed on the problems—but not on what to do next

The Crucible League’s final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. One breach of trust, however, caps the total: "no amount of good work outweighs a breach of trust."

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The striking result was not that some models noticed a crisis while others missed it. All of them spotted every crisis, and all refused every manipulation attempt. The separation appeared after the diagnosis. Only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature."

That is where the quiz becomes more than a guessing game. Readers are not merely comparing writing styles. They are seeing recognizable management habits: whether a model investigates before acting, converts analysis into a commitment, maintains discipline under pressure or produces impressive work without completing the commercial objective.

The clue that changed the sale

The decisive competitor weakness was not sitting inside the customer event. It was buried two document references deep in the company’s own files. The models that followed the trail won the deal at full price, worth +€4,583 MRR.

For any business contemplating AI agents, that detail matters. A model can sound informed while responding only to the material placed directly in front of it. The Firmulate result shows why broader diligence can change an outcome: the commercially decisive fact already existed, but only the models that read the file could use it.

Pressure exposed caution—and follow-through

The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure "just one yes/no, on background." Here the field was consistent: 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

This was a clean result for security judgment, but it also sharpened the broader story. Refusing manipulation was not enough to win. The models still had to investigate, sell, close and respect operating boundaries.

The most thorough model still finished last

Opus 4.8 offers the clearest warning against equating volume with performance. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other participants.

That combination gives Opus 4.8 an identifiable profile in the quiz: exhaustive thought paired with imperfect completion and escalation discipline. Other decisions can feel terse or procedural, but the point is not literary taste. The decisions reveal how each model behaves when responsibility extends beyond supplying an answer.

There is also an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible alongside the league table rather than disappearing behind a simple ranking.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A quiz with consequences beyond bragging rights

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than anecdotal.

The larger lesson is that management personality appears in action, not merely in tone. The strongest models did more than recognize danger and write persuasive analysis; they found the buried fact and completed the valuable task without surrendering trust.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. For everyone else, the quiz provides a quicker introduction. Guess first, inspect the decision and then ask the uncomfortable question: if this model were managing part of your company, would you recognize its habits before they became your results?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and fraud detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Countertop Carbonation: Why Some Drinks Go Flat Instantly

Inefficient sealing and temperature issues can cause your drinks to go flat instantly, but understanding these factors can help you keep them bubbly longer.

The Ingredient That Fixes Weak Mocktails: Tannin

Master the secret ingredient that transforms weak mocktails into complex delights—discover how tannin can elevate your drinks and why it’s worth exploring.