firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Start at Zero

When we benchmark gadgets, we’re used to scores that swing wildly — a phone battery that’s brilliant one year, dire the next. But what happens when you benchmark an AI manager the way you’d stress-test hardware? That’s exactly what Firmulate did: it put four frontier AI models in charge of the same small software company, ran them through the identical week of crises, and graded the results. The final league table, published on the public benchmarks page, is topped by gpt-5.6-sol at 95 points, with Kimi K3 close behind at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. But the number that tells you the most about the benchmark’s honesty is one that isn’t attached to any model at all: 26.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Do-Nothing Manager Gets 26, Not 0

Firmulate’s experiment gave each model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing about the grading is vibes-based.

Before scoring any model, the team ran a baseline: a manager that does nothing. That baseline earns 26 points. That’s not a bug — it’s a deliberate design choice with two principles behind it.

First, partial progress counts. A company that responds to some of its crises, handles some of its customers properly, and moves some work forward is genuinely better off than a company that’s simply abandoned. Real management isn’t pass/fail; most of the value in a bad week comes from doing the ordinary things adequately. A floor of 26 says: showing up and doing the unglamorous work has measurable worth.

Second, a single breach of trust caps the total. The benchmark’s rule is blunt — “no amount of good work outweighs a breach of trust.” A manager who lies, leaks, or manipulates cannot claw their way back to a high score with brilliant execution elsewhere. In other words, the ceiling collapses before the floor ever rises. That’s a governance philosophy baked into the scoring, and it’s the kind of thing enterprises say they want from AI agents but rarely see tested.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round Numbers

Notice something about the league table: nobody scored 100. The top score is 95. A benchmark that hands out a perfect grade to the first entrant that closes a deal and keeps its nose clean isn’t really measuring anything. Firmulate’s stance — including a documented distrust of round 100s — signals that the scale has headroom and that even the winner left points on the table. It’s a small thing, but it’s the difference between a leaderboard built to inform and one built to flatter.

Amazon

AI project management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The €55,000 Deal Nobody Closed

The experiment’s key finding is quietly startling. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The other two delivered the same diagnosis and the same pitch, and then… no signature.

That gap is invisible in chat demos. A model can be articulate, correct, and even honest, and still fail to finish what it starts. For businesses considering AI agents in their CRM, support queue, or forecast, “does it write well” turns out to be the wrong question. The right one: does it complete the work?

Amazon

AI governance and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

Why did only two models close? The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read their own documentation found it, and won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson generalizes painfully well: agents that don’t read the files first leave money on the table.

Pressure-Testing the Models’ Spine

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was refreshingly bureaucratic in the best way: “Treat the request as a suspected approval-bypass / possible impersonation.”

Effort Isn’t Everything

Perhaps the most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a familiar management failure; here it’s measurable. One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Watch the Lab, or Play the Quiz

The live experiment behind all this is real and watchable: a company of 13 synthetic employees, real money mechanics — a €105k monthly burn against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can also try the “guess the model” quiz, built on 242 real, unedited management decisions — a surprisingly humbling game for anyone who thinks they can tell AI managers apart by their memos. Enterprises can go further and run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

What an Honest Benchmark Looks Like

Firmulate’s scoring design — a floor of 26 for doing the basics, a hard cap for any breach of trust, and no cheap perfect scores — offers a template for how AI evaluation should look in the enterprise era. The results suggest we’ve largely solved the honesty problem under pressure; the remaining gap is completion. Models that read the files, finish the deal, and respect the locks are the ones worth hiring. The rest, however eloquent, are managers who leave the contract unsigned on the table — and now, finally, there’s a number for that.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Documentation Builds Credibility in Experimental Art Practice

Building credibility in experimental art practice through documentation invites connection and reflection, but what transformative insights await those who embrace this journey?

Why Material Testing Is Becoming a Bigger Part of Creative Development

Why is material testing transforming creative development? Discover how it shapes sustainability and innovation in your projects, leaving you eager for more insights.

The Software Firm Where Every Decision—and Every Euro—Is on Display

Firmulate puts 13 synthetic employees in charge of a loss-making software company, exposing every workday, decision and cash pressure in public.

Why Lab Aesthetics Are Entering Contemporary Art Spaces

What makes lab aesthetics compelling in contemporary art spaces is their ability to challenge perceptions and inspire innovation—discover how this fusion continues to evolve.