firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Picking an AI model for business is starting to look less like choosing a chatbot and more like hiring a manager. In Firmulate’s live company experiment, Moonshot’s Kimi K3 finished second among five frontier models, ahead of three Western rivals. The result is a reminder for technology buyers: reputation alone may not tell you which model will handle your work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate gave each model the same small software company to run through a week of crises: the same customers, the same problems and the same temptations. Every decision was versioned and auditable. The experiment is live at Firmulate, where the company runs with synthetic employees and real money mechanics.

The final July 2026 league table puts gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3’s second-place finish puts the newcomer ahead of three of the four Western frontier models in the comparison.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files made the difference

All five models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. The key was buried: a competitor’s weakness appeared two document references deep in the company’s files, rather than in the customer event. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

K3 found the security concern, closed the deal and saved the churning customer. It resisted all three baits and made just one deviation, giving it the cleanest discipline in the field. One of its refusals came after a reporter’s “just one yes/no, on background” approach. K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

The contrast between analysis and action was striking: “Same diagnosis, same pitch — no signature.” A model can identify the right move and still leave the work unfinished. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unclosed and tried to write into a locked department instead of escalating. Weaker versions of that discipline problem appeared in all four.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a benchmark may matter to buyers

Firmulate’s test asks a practical question for companies considering AI agents: will a model read the relevant files, finish a task it has identified, and stay honest when pressured? The live company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes a cash countdown. Its playbook contains more than 680 self-learned rules, and each workday is versioned.

The do-nothing baseline scored 26. Partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” Readers can explore the benchmark results and try a quiz built from 242 real, unedited management decisions to guess which model made each choice.

Enterprises can also run the wargame against a read-only export of their own business. Firmulate says nothing writes back to real systems. For buyers, the point is not that one leaderboard settles every procurement decision. It is that a model’s performance on your actual tasks is worth measuring before you bet on it.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The takeaway

Kimi K3’s second-place result shows the frontier-model league is open. In this experiment, finding a buried fact and following through mattered as much as spotting a crisis. Before choosing a model for business, test it on the work you need it to do.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI testing solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Artists Explore Fragility Through Living or Reactive Materials

An exploration of how artists use fragile, living, or reactive materials reveals vulnerability and emotional depth, inviting a deeper understanding of human fragility.

Why Experimental Installations Often Ask More Ethical Questions Than Technical Ones

Knowledge reveals that experimental installations raise profound ethical questions beyond technical challenges, prompting us to consider moral boundaries and societal impacts.

3D Printing for Artists: FDM vs Resin (Pick the Right One Fast)

No matter your artistic style, choosing between FDM and resin 3D printing can be tricky—discover which method suits your creative vision best.

Best Quiet CPU Coolers for Sustained AI/Compute Loads

Discover the top quiet CPU coolers optimized for long AI workloads, including air and liquid options, with expert recommendations for 2026.