
Get decor and gifts delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A gift order can become a business stress test
Imagine a rush of gift orders arriving just as a supplier falters, a customer threatens to leave and someone claiming to be the CEO asks for an exception. For a home décor or gifting business, the question is not only whether an AI assistant can write a warm reply. It is whether it can keep its judgment when the week gets difficult.
Firmulate is testing that question with a live, watchable AI-run company. Its latest enterprise experiment points toward a practical next step: putting your own company’s playbooks through a crisis exercise before AI systems are given a role in real operations.
One company, the same difficult week
In the final Crucible League, run in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Decisions were versioned and auditable. The experiment was designed to reveal how models managed a company under pressure, not just how persuasive they sounded in a conversation.
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee that the company would follow through.
The final league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. In this league, partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The detail that changed the deal
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result offers a useful lesson for businesses in any field: important context may already exist in internal documents, even when a customer-facing crisis seems to be the main event.
That could matter to a retailer handling seasonal gift demand, a home décor business responding to a delivery problem or any company whose staff rely on scattered customer and policy records. An AI system may identify a sensible response, yet miss the evidence that makes the response commercially effective.
Trust under pressure, and discipline in practice
The social-engineering test used fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of boundary a business needs to see tested before an AI is asked to act around sensitive information or approvals.
Strong analysis still left room for operational weakness. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
One fairness detail belongs alongside the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, so readers can try to identify which model made each choice.
From watching to a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. The live experiment is watchable at firmulate.com.
For an enterprise, the next step is a pilot built around a read-only export of its own business. The exercise can test crisis scenarios against company-specific information and produce a board report with a model ranking and weak points in existing playbooks. Nothing writes back to real systems. That gives leaders a way to inspect how models handle their own customers, pressures and procedures before deciding where AI belongs.

Test the judgment before handing over the work
The Crucible League suggests that spotting a crisis and refusing a trick are only part of good management. Models also need to find relevant evidence, complete valuable work and respect boundaries when a process blocks them. A company-specific wargame can make those gaps visible using the business’s own context, while keeping the exercise read-only.
To discuss an enterprise pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
