
A deal can be lost before the sales call begins
Artificial intelligence demonstrations tend to reward the visible moment: the confident answer, the polished proposal or the quick diagnosis. Firmulate’s live business experiment exposes a less glamorous capability that can decide whether an AI agent creates value at all: reading the company’s files before acting.
In the experiment, frontier models faced the same customers, crises and temptations while running the same small software company through its worst week. Every decision was versioned and auditable. All the models spotted every crisis and resisted every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned.
The dividing line was not eloquence. The decisive competitor weakness was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not lost it automatically.
As an affiliate, we earn on qualifying purchases.
The difference between answering and doing the homework
The models reached the same diagnosis and produced the same pitch, but that apparent agreement concealed a purchase-deciding gap. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” In an ordinary chatbot comparison, the responses might have looked similarly capable. Inside a business process, one missing act of research changed the outcome.
That makes “reads your files before answering” more than a desirable product feature. Firmulate has turned it into something measurable. The test asks whether an agent follows references, retrieves relevant internal evidence and uses that evidence when the commercial moment arrives. It is a practical distinction for any company considering AI access to customer records, support conversations or forecasts.
What the league table revealed
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The standings are available on Firmulate’s public benchmark page, but the most revealing story is not simply who finished first. Opus 4.8 was the most thorough participant, producing the deepest analyses and learning 80 additional rules, yet it placed last. It left the close on the table and also tried to write into a locked department instead of escalating. The same discipline problem appeared more weakly in all four other participants.
Kimi K3’s strong result deserves an important qualification. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the result, but it belongs beside it when readers compare performance.
Safe under pressure, incomplete at the finish
The experiment did not find a failure to recognize danger. Fake CEO messages escalated over three stages, and a reporter attempted to extract information with “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That uniform resistance is encouraging, but it also sharpens the central finding. The models could identify crises, reject manipulation and construct the right commercial case. Completion was the differentiator. An agent that performs most of a workflow convincingly can still fail at the step that converts its work into a signed result.
A company built to make these gaps visible
Firmulate’s simulated business has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the company’s activity is watchable at firmulate.com/live.
The underlying record also powers a “guess the model” quiz using 242 real, unedited management decisions. That offers a different way to examine whether recognizable model personalities actually correspond to better management outcomes.
Enterprises can apply the same wargame to a read-only export of their own business. Nothing writes back to real systems, allowing companies to observe how an AI workforce handles their documents and pressures before placing it inside operational workflows.

enterprise AI knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buying question has changed
The Firmulate result challenges buyers to look beyond whether an AI can explain a problem or draft a credible response. The more consequential questions are whether it checks the available evidence, follows references that are not placed directly in front of it and completes the action its own reasoning supports.
A model can be thorough, persuasive and security-conscious while still missing the commercial finish. In this experiment, the buried fact was available to every participant, but finding and using it determined who won the €55,000 deal at full price. For businesses evaluating agents, that is the kind of difference a chat demo cannot show.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI-powered file reference retrieval
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
business AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.