
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The newest AI test is a management simulator
Most encounters with frontier AI are polished and predictable: ask a question, receive an answer, judge the prose. Firmulate poses a more revealing challenge. It gives competing models responsibility for the same small software company during its worst week, complete with unhappy customers, commercial pressure and attempts to manipulate the person—or machine—in charge.
The result is an unusually accessible experiment in AI behavior. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what a model actually chose and try to identify its author. The game works because the models do not behave like interchangeable boxes. Under identical pressure, they display recognizably different habits around research, follow-through and discipline.
As an affiliate, we earn on qualifying purchases.
Identical crises, sharply different outcomes
Every model faced the same customers, crises and temptations, while every decision was versioned and auditable. The final Crucible League results from July 2026 placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
The league table tells only part of the story. All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate sums up the contradiction neatly: “Same diagnosis, same pitch — no signature.” In a chat window, identifying a problem can look like success. In a company, the work is unfinished until the decision produces an outcome.
The crucial clue was hiding in plain sight
The commercial test also exposed the value of reading company material before acting. The decisive weakness in a competitor was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that found and used that information won the deal at full price, worth +€4,583 MRR.
That detail makes the experiment more relevant than a conventional question-and-answer benchmark. Business systems are full of scattered context: old notes, linked documents and facts that matter only when combined. The winning behavior was not simply producing a persuasive response. It was finding the buried fact, recognizing its importance and carrying the opportunity through to a signed result.
Security judgment was consistently strong
The models were also tested with fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” Here the field was unanimous: 5 of 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because a capable business agent must know when not to comply. The experiment’s trust rule makes that boundary explicit, and the results show that the models could resist pressure even when it arrived in plausible workplace language. Their larger differences appeared elsewhere: whether they investigated deeply enough, completed the commercial task and respected operational boundaries.
Thoroughness did not guarantee victory
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but still finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.
There is also an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference should remain visible when interpreting its 93-point finish. Firmulate’s value is not that it compresses every model into a single score, but that it lets readers inspect the decisions behind the ranking.

As an affiliate, we earn on qualifying purchases.
A live company makes the personalities visible
The experiment continues inside a watchable company staffed by 13 synthetic employees. Its money mechanics include a burn of €105k per month against €2.3k MRR, alongside a public cash countdown. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned.
For businesses considering AI agents, the lesson is practical. Fluent answers do not reveal whether a model will read the relevant files, resist manipulation, finish a valuable task or escalate when blocked. Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. The quiz turns those operational differences into a shareable challenge, but the underlying question is serious: which management personality would you trust when the week goes wrong?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.