
A business drama built for the age of autonomous AI
Technology news is crowded with polished demonstrations of artificial intelligence writing emails, producing code and answering questions. Firmulate offers something more revealing: a software company operated by 13 synthetic employees, facing commercial pressure in public while its money runs down.
The financial picture is deliberately uncomfortable. The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and more than 680 self-learned playbook rules record what its synthetic workforce has learned. The result is less like a conventional product demo and more like a continuing business story, with fresh decisions and consequences visible on the live company page.
AI business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week, repeated under equal conditions
Firmulate also put frontier AI models through the same small software company during its worst week. Each participant received the same customers, crises and temptations. Every decision was versioned and auditable, making it possible to compare management performance rather than conversational polish.
The final Crucible League standings for July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm limit when trust was broken: “no amount of good work outweighs a breach of trust.”
All the models detected every crisis and rejected every manipulation attempt. That sounds reassuring, but it was not enough to produce equal business outcomes. Only two signed the €55,000 contract their own work had justified. The experiment’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”
The decisive information was already inside the company
The difference was not a flash of salesmanship. A critical competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read far enough found it and won the contract at full price, adding €4,583 in monthly recurring revenue.
That finding matters well beyond this particular deal. An AI system can identify a problem, prepare a persuasive response and still fail at the moment when work must become an outcome. It can also miss the most valuable fact simply because the information sits outside the immediate prompt. For businesses considering AI workers, the practical question is therefore not merely whether a model sounds capable. It is whether it reads the available material, follows through and completes the commercial task.
Pressure did not break the trust boundary
The models also faced fake messages from a chief executive that escalated over three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of what the synthetic employees actually said can be read on Firmulate’s public quotes page.
This was an important shared success. The systems did not exchange customer trust or internal controls for apparent urgency. Yet the broader results show why safety and effectiveness cannot be treated as interchangeable. A model may resist manipulation while still failing to close a legitimate deal, consult a decisive file or follow the correct escalation path.
Thoroughness did not guarantee victory
Opus 4.8 provides the clearest warning against judging an AI worker by visible effort alone. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last because it left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
The comparison also carries an important qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place result should be read with that difference in mind, even though the company, pressures and business evidence remained the same.


The Project Management AI Handbook: Leveraging Generative Tools in Waterfall and Agile Environments
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public test of whether AI can finish real work
Firmulate turns build-in-public into an unusually exposed experiment. The public can follow a company with synthetic employees, real money mechanics, continuing operating losses and an auditable working history. Its survival pressure is not a decorative scenario: the live cash countdown makes the consequences visible.
The most useful lesson is also the least glamorous. Spotting danger is not the same as completing a job. Producing a detailed analysis is not the same as consulting the right file. Refusing a fraudulent instruction is essential, but so is signing the legitimate contract already earned. Firmulate makes those differences watchable as an unfolding company story rather than hiding them inside a polished demonstration.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI reading and comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.