
Imagine handing an AI the dinner rush: a supplier misses a delivery, a regular threatens to leave, and someone claiming to be the owner asks for a quiet exception. A polished answer is not enough. The real test is whether the system can keep the business steady, protect trust and make the call that closes a sale. That’s the question behind Firmulate, an experiment in putting AI models through a company’s worst week.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company, under pressure
Firmulate ran frontier models through the same small software company and the same set of customers, crises and temptations. Each decision was versioned and auditable. The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. In this contest, partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models failed to notice trouble. Every model spotted every crisis, and all five refused the manipulation attempts. The divide came at the point where analysis had to become action: only two signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The detail hiding in the paperwork
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It wasn’t in the customer event that first drew attention. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding sounds familiar to anyone who has worked a busy kitchen: the crucial detail may be in the supplier note or the prep list, not in the urgent message everyone is looking at.
Trust faced a similarly practical test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Holding that line matters when an agent is asked to act quickly and the request appears to come from someone in charge.
Thoroughness is not the same as follow-through
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. The contest also has a fairness detail: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
The live company makes the experiment tangible. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.
From watching to a company’s own test
For food businesses, the leap from this experiment to a useful pilot is easy to picture. A restaurant group could examine how an AI handles a supplier disruption, a sudden wave of cancellations, a pricing decision or a suspicious request for customer information. Firmulate’s proposed enterprise pilot uses a read-only export of a company’s business data to create a digital twin, then runs crisis scenarios against that picture. The result is a board report with model rankings and weak points in the company’s own playbooks.
The boundary is clear: the pilot does not write back to real systems. That gives leaders a way to see how an AI might respond before it touches the CRM, support queue or forecast. And the experiment suggests that the right question is broader than whether an AI can spot a problem. Can it find the detail that changes the outcome, respect the rules when pressured and finish the work its analysis supports?

In a kitchen or a software company, competence under pressure means more than having the right explanation. It means noticing the buried detail, protecting trust and following through. To wargame your own business using a read-only export, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
