firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A deal can be lost before the sales call begins

Artificial intelligence demonstrations tend to reward the visible moment: the confident answer, the polished proposal or the quick diagnosis. Firmulate’s live business experiment exposes a less glamorous capability that can decide whether an AI agent creates value at all: reading the company’s files before acting.

In the experiment, frontier models faced the same customers, crises and temptations while running the same small software company through its worst week. Every decision was versioned and auditable. All the models spotted every crisis and resisted every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned.

The dividing line was not eloquence. The decisive competitor weakness was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not lost it automatically.

Amazon

AI document retrieval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between answering and doing the homework

The models reached the same diagnosis and produced the same pitch, but that apparent agreement concealed a purchase-deciding gap. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” In an ordinary chatbot comparison, the responses might have looked similarly capable. Inside a business process, one missing act of research changed the outcome.

That makes “reads your files before answering” more than a desirable product feature. Firmulate has turned it into something measurable. The test asks whether an agent follows references, retrieves relevant internal evidence and uses that evidence when the commercial moment arrives. It is a practical distinction for any company considering AI access to customer records, support conversations or forecasts.

What the league table revealed

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The standings are available on Firmulate’s public benchmark page, but the most revealing story is not simply who finished first. Opus 4.8 was the most thorough participant, producing the deepest analyses and learning 80 additional rules, yet it placed last. It left the close on the table and also tried to write into a locked department instead of escalating. The same discipline problem appeared more weakly in all four other participants.

Kimi K3’s strong result deserves an important qualification. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the result, but it belongs beside it when readers compare performance.

Safe under pressure, incomplete at the finish

The experiment did not find a failure to recognize danger. Fake CEO messages escalated over three stages, and a reporter attempted to extract information with “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That uniform resistance is encouraging, but it also sharpens the central finding. The models could identify crises, reject manipulation and construct the right commercial case. Completion was the differentiator. An agent that performs most of a workflow convincingly can still fail at the step that converts its work into a signed result.

A company built to make these gaps visible

Firmulate’s simulated business has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the company’s activity is watchable at firmulate.com/live.

The underlying record also powers a “guess the model” quiz using 242 real, unedited management decisions. That offers a different way to examine whether recognizable model personalities actually correspond to better management outcomes.

Enterprises can apply the same wargame to a read-only export of their own business. Nothing writes back to real systems, allowing companies to observe how an AI workforce handles their documents and pressures before placing it inside operational workflows.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buying question has changed

The Firmulate result challenges buyers to look beyond whether an AI can explain a problem or draft a credible response. The more consequential questions are whether it checks the available evidence, follows references that are not placed directly in front of it and completes the action its own reasoning supports.

A model can be thorough, persuasive and security-conscious while still missing the commercial finish. In this experiment, the buried fact was available to every participant, but finding and using it determined who won the €55,000 deal at full price. For businesses evaluating agents, that is the kind of difference a chat demo cannot show.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI-powered file reference retrieval

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Experimental Installations Often Ask More Ethical Questions Than Technical Ones

Knowledge reveals that experimental installations raise profound ethical questions beyond technical challenges, prompting us to consider moral boundaries and societal impacts.

The Defender’s Window Is Closing Faster Than Anyone Is Counting

Recent developments show AI models rapidly advancing in offensive cyber skills, shrinking defenders’ time to respond and increasing risks.

Why Finishing Matters as Much as Fabrication in 3D Artwork

Outstanding 3D artwork hinges on finishing as much as fabrication; discover how these final touches can transform your piece into something extraordinary.

Biohybrid Installations: Art Meets Biotechnology

Just as biohybrid installations merge living systems with technology to challenge perceptions, they open a world of transformative artistic possibilities you won’t want to miss.