
A coffee company lives on more than a good blend. It depends on suppliers, loyal customers, reliable delivery and the judgment to protect trust when pressure hits. As AI takes on more business decisions, owners in coffee and beyond face a practical question: how would an AI workforce handle a bad week inside their company?
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
One company, one difficult week
Firmulate’s Crucible League put frontier AI models in charge of the same small software company through its worst week. Each faced the same customers, crises and temptations. The experiment tracked every decision, making the results auditable.
In the final league, dated July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s stated principle is that partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Seeing the problem was not the same as acting
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s concise description of that gap: “Same diagnosis, same pitch — no signature.”
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a useful distinction for any business considering AI: an agent might recognize the issue and produce a convincing recommendation, but still fail to complete the work—or miss the detail that changes the outcome.
Trust faced a direct test
The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its judgment on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters to companies where customer data, pricing or internal plans could be exposed by a seemingly casual request. But the league also surfaced a different kind of weakness: Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
A live company, and a possible next step
Firmulate’s live company is synthetic, with 13 employees and real money mechanics: burn of €105k a month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
For an enterprise, the proposed next step is a pilot using a read-only export of its own business. Crisis scenarios can be run against that digital twin, producing a board report with model rankings and weak points in the company’s playbooks. The pilot does not write back to real systems. That creates a way to examine how AI handles a company’s customers, rules and pressure points before it is trusted with live work.

Try the wargame against your business
To explore a Firmulate enterprise pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
