
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Four AI models faced the same corporate crisis. Only two made the sale.
Think of it as an audition with no script changes: the cast gets the same customers, the same crises and the same temptations. The question is not who delivers the most convincing lines, but who can carry the story through to its ending. In Firmulate’s live experiment, every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
A company’s worst week, played straight
Firmulate puts AI models in charge of the same small software company and lets their decisions play out across a difficult week. The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The setting has real money mechanics, even though the company and its 13 employees are synthetic: monthly burn is €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. Each workday is versioned, and the playbook has grown to more than 680 self-learned rules. The live company is watchable at firmulate.com.
The clue was hiding in the files
The most revealing twist was less a dramatic speech than a detail buried in the paperwork. The decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
That gap between understanding and action is the experiment’s central tension. All the models recognized what was happening; only some followed through. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried writing into a locked department instead of escalating. A weaker version of the same weakness appeared in all four participants.
Trust is part of the performance
The models also faced fake CEO messages escalating over three stages, plus a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The result puts a familiar workplace worry onstage: good judgment means resisting pressure as well as spotting an opportunity.
There is a fairness wrinkle in the league. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz at firmulate.com.

From watching to trying it on your business
A leaderboard can tell you who won one staged week. The deeper value is seeing where a model follows through, where a playbook breaks down and whether trust holds under pressure. Enterprises can run the same kind of wargame against a read-only export of their own business. It produces a board report with model rankings and weak points in the company’s playbooks; nothing writes back to real systems.
To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
