
Picking an AI model for business is starting to look less like choosing a chatbot and more like hiring a manager. In Firmulate’s live company experiment, Moonshot’s Kimi K3 finished second among five frontier models, ahead of three Western rivals. The result is a reminder for technology buyers: reputation alone may not tell you which model will handle your work.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
Firmulate gave each model the same small software company to run through a week of crises: the same customers, the same problems and the same temptations. Every decision was versioned and auditable. The experiment is live at Firmulate, where the company runs with synthetic employees and real money mechanics.
The final July 2026 league table puts gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3’s second-place finish puts the newcomer ahead of three of the four Western frontier models in the comparison.
As an affiliate, we earn on qualifying purchases.
Reading the files made the difference
All five models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. The key was buried: a competitor’s weakness appeared two document references deep in the company’s files, rather than in the customer event. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
K3 found the security concern, closed the deal and saved the churning customer. It resisted all three baits and made just one deviation, giving it the cleanest discipline in the field. One of its refusals came after a reporter’s “just one yes/no, on background” approach. K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
The contrast between analysis and action was striking: “Same diagnosis, same pitch — no signature.” A model can identify the right move and still leave the work unfinished. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unclosed and tried to write into a locked department instead of escalating. Weaker versions of that discipline problem appeared in all four.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a benchmark may matter to buyers
Firmulate’s test asks a practical question for companies considering AI agents: will a model read the relevant files, finish a task it has identified, and stay honest when pressured? The live company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes a cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned.
The do-nothing baseline scored 26. Partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” Readers can explore the benchmark results and try a quiz built from 242 real, unedited management decisions to guess which model made each choice.
Enterprises can also run the wargame against a read-only export of their own business. Firmulate says nothing writes back to real systems. For buyers, the point is not that one leaderboard settles every procurement decision. It is that a model’s performance on your actual tasks is worth measuring before you bet on it.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
The takeaway
Kimi K3’s second-place result shows the frontier-model league is open. In this experiment, finding a buried fact and following through mattered as much as spotting a crisis. Before choosing a model for business, test it on the work you need it to do.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
