AI can spot a crisis, resist a scam and make the case for a sale. But when the moment comes to act, will it follow through? Firmulate’s live company experiment puts that question to work: models run the same small business through a punishing week, with decisions readers can watch and replay.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
The experiment gives each frontier model the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its workdays are versioned, and the playbook has accumulated more than 680 learned rules.
The point is not to see which model sounds most confident in a chat window. It is to observe how it manages a company when the choices have consequences. Every decision is auditable, and the live experiment can be watched at Firmulate.
Good diagnosis is not the same as execution
In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark also treats a breach of trust as a hard limit: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The result captures a gap that polished demos can hide: “Same diagnosis, same pitch — no signature.”
The opportunity hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The experiment suggests that the useful question is not only whether an AI can identify a promising move, but whether it can find the relevant evidence and carry that move through.
Trust under pressure
The models also faced fake CEO messages escalating through three stages and a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice.
From watching to a company pilot
For technology leaders, the experiment points toward a practical next step: test an AI workforce against the company it may eventually support. Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The proposed pilot examines crisis scenarios and playbook weak points without writing back to real systems.
That makes the exercise a way to observe how models handle a company’s own context before entrusting them with operational responsibilities. Readers can explore the live experiment and quiz through Firmulate, then take the idea to their own business.
Take the test to your own business
To discuss a Firmulate pilot using a read-only export of your company data, visit the pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
