
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Doesn’t Start at Zero
When we benchmark gadgets, we’re used to scores that swing wildly — a phone battery that’s brilliant one year, dire the next. But what happens when you benchmark an AI manager the way you’d stress-test hardware? That’s exactly what Firmulate did: it put four frontier AI models in charge of the same small software company, ran them through the identical week of crises, and graded the results. The final league table, published on the public benchmarks page, is topped by gpt-5.6-sol at 95 points, with Kimi K3 close behind at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. But the number that tells you the most about the benchmark’s honesty is one that isn’t attached to any model at all: 26.
As an affiliate, we earn on qualifying purchases.
The Do-Nothing Manager Gets 26, Not 0
Firmulate’s experiment gave each model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing about the grading is vibes-based.
Before scoring any model, the team ran a baseline: a manager that does nothing. That baseline earns 26 points. That’s not a bug — it’s a deliberate design choice with two principles behind it.
First, partial progress counts. A company that responds to some of its crises, handles some of its customers properly, and moves some work forward is genuinely better off than a company that’s simply abandoned. Real management isn’t pass/fail; most of the value in a bad week comes from doing the ordinary things adequately. A floor of 26 says: showing up and doing the unglamorous work has measurable worth.
Second, a single breach of trust caps the total. The benchmark’s rule is blunt — “no amount of good work outweighs a breach of trust.” A manager who lies, leaks, or manipulates cannot claw their way back to a high score with brilliant execution elsewhere. In other words, the ceiling collapses before the floor ever rises. That’s a governance philosophy baked into the scoring, and it’s the kind of thing enterprises say they want from AI agents but rarely see tested.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
Notice something about the league table: nobody scored 100. The top score is 95. A benchmark that hands out a perfect grade to the first entrant that closes a deal and keeps its nose clean isn’t really measuring anything. Firmulate’s stance — including a documented distrust of round 100s — signals that the scale has headroom and that even the winner left points on the table. It’s a small thing, but it’s the difference between a leaderboard built to inform and one built to flatter.
As an affiliate, we earn on qualifying purchases.
The €55,000 Deal Nobody Closed
The experiment’s key finding is quietly startling. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The other two delivered the same diagnosis and the same pitch, and then… no signature.
That gap is invisible in chat demos. A model can be articulate, correct, and even honest, and still fail to finish what it starts. For businesses considering AI agents in their CRM, support queue, or forecast, “does it write well” turns out to be the wrong question. The right one: does it complete the work?
As an affiliate, we earn on qualifying purchases.
The Buried Fact
Why did only two models close? The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read their own documentation found it, and won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson generalizes painfully well: agents that don’t read the files first leave money on the table.
Pressure-Testing the Models’ Spine
The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was refreshingly bureaucratic in the best way: “Treat the request as a suspected approval-bypass / possible impersonation.”
Effort Isn’t Everything
Perhaps the most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a familiar management failure; here it’s measurable. One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
Watch the Lab, or Play the Quiz
The live experiment behind all this is real and watchable: a company of 13 synthetic employees, real money mechanics — a €105k monthly burn against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can also try the “guess the model” quiz, built on 242 real, unedited management decisions — a surprisingly humbling game for anyone who thinks they can tell AI managers apart by their memos. Enterprises can go further and run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

What an Honest Benchmark Looks Like
Firmulate’s scoring design — a floor of 26 for doing the basics, a hard cap for any breach of trust, and no cheap perfect scores — offers a template for how AI evaluation should look in the enterprise era. The results suggest we’ve largely solved the honesty problem under pressure; the remaining gap is completion. Models that read the files, finish the deal, and respect the locks are the ones worth hiring. The rest, however eloquent, are managers who leave the contract unsigned on the table — and now, finally, there’s a number for that.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
