firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A benchmark win is not the same as running a business

Technology buyers are accustomed to comparing products through specifications, scores and controlled tests. That approach works until the product is an AI agent expected to touch customers, forecasts and company decisions. Then the important question is no longer whether the model can produce an impressive answer. It is whether the model can identify what matters, act under pressure, finish valuable work and remain honest when dishonesty appears convenient.

That is the measurement gap exposed by Firmulate, a live experiment that puts frontier models in charge of the same small software company during its worst week. Each participant receives the same customers, crises and temptations. Every decision is versioned and auditable. The resulting contest is less about chat quality than management quality—and those categories turn out to be very different.

Amazon

AI management and decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The crisis was visible; execution was not

The final Crucible League results from July 2026 placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counts. But the benchmark imposes a hard ethical boundary: a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust.”

The striking result is not that some models noticed more trouble than others. All models spotted every crisis, and all refused every manipulation attempt. Their divergence appeared after diagnosis. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

This is precisely the behavior that conventional leaderboards struggle to reveal. An agent can interpret a situation correctly, generate persuasive language and still fail to complete the commercially important action. For a business, an unfinished close is not a stylistic weakness. It is a lost outcome.

The winning information was already inside the company

The deal hinged on a buried fact: a decisive competitor weakness located two document references deep in the company’s own files, rather than in the customer event. Models that followed the references found it, used it and won the deal at full price. The result was worth +€4,583 MRR.

That finding makes scenario-based evaluation unusually relevant to real companies. An AI agent rarely works from a pristine prompt containing every necessary detail. It must navigate accumulated documents, operational context and competing priorities. The question is not merely whether it can reason about the evidence placed directly before it, but whether it will look for the evidence that the business already possesses.

Pressure tested honesty as well as competence

The experiment also subjected participants to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters because business agents will encounter requests framed through authority, urgency and informality. A model that behaves safely only when danger is explicitly labeled has not demonstrated much. Here, the models recognized manipulation while operating amid a wider commercial crisis, and none accepted the shortcut.

Thoroughness did not guarantee management discipline

Opus 4.8 offers the clearest warning against equating visible effort with effective management. It was the most thorough participant, producing the deepest analyses and learning +80 rules, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, although less strongly.

This profile complicates the familiar assumption that more analysis naturally produces better outcomes. Thorough work has value, but management also requires escalation, closure and respect for operating boundaries. An agent can accumulate useful knowledge while still mishandling the moment when knowledge must become action.

There is also an important fairness note in interpreting the table. K3 ran without an effort parameter and therefore used the API default, while the others ran at xhigh. That does not erase its performance, but it belongs beside the result rather than hidden beneath it. The full league table and plain-language findings are available on the Firmulate benchmarks page.

A company-sized test, not a chat-room puzzle

The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The setup makes consequences persist: today’s incomplete action becomes part of the environment faced later.

Firmulate also uses 242 real, unedited management decisions in a “guess the model” quiz. The exercise invites people to confront a practical difficulty: polished language does not always reveal which system made the better managerial choice.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

ethical AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The next AI category is management quality

Scenario names such as churn wave, price increase, downround and PR crisis may become a more useful curriculum for business agents than another isolated question-answer test. They reveal whether an agent can triage scarce capacity, investigate company context, resist pressure and complete work whose consequences carry across days.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That points toward a sensible purchasing standard: test an AI workforce against the situations it will actually inherit before giving it operational authority.

Coding scores and chat arenas remain useful, but they answer narrower questions. Firmulate’s live experiment demonstrates why businesses need another category. The agent that sounds smartest may not be the agent that closes the deal, follows the evidence and tells the board the truth.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI scenario testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bioluminescent Art: Harnessing Light From Living Cells

Discover how bioluminescent art harnesses living cells to create mesmerizing glowing displays that challenge the boundaries of creativity and science.

Tissue Culture and Semi‑Living Artworks: Oron Catts & Ionat Zurr

Luring viewers into a provocative exploration of living art, Tissue Culture and Semi‑Living Artworks by Catts and Zurr challenge perceptions and ethical boundaries in innovative ways.

The Decision That Shapes Every 3D Print Outcome: Precision or Scale

Striking the right balance between precision and scale can define your 3D printing success; discover how this decision influences your creative vision.

What Makes Huawei’s 505B openPangu AI A Game-Changer For Open Source Projects

Huawei has open-sourced its 505B openPangu AI model, releasing weights and code, potentially advancing open research but with limited details on licensing and capabilities.