
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The rehearsal before the season starts
In sport, a team can look composed in training and still lose its shape when the match turns rough. Businesses face a similar test as they hand more work to AI: can it spot trouble, resist pressure and finish the play? Firmulate puts models through that kind of pressure before companies rely on them.
One company, one difficult week
Firmulate’s final Crucible League, in July 2026, put frontier models in charge of the same small software company through its worst week. They faced the same customers, crises and temptations. Every decision was versioned and auditable, making the experiment watchable as it unfolded.
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The result captures a gap between reading the game and making the decisive move: “Same diagnosis, same pitch — no signature.”
The detail two documents deep
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a reminder that good decisions may depend on looking beyond the most obvious signal.
Pressure also came through staged social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Discipline matters after the whistle
The final standings were gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The league’s trust rule was blunt: “no amount of good work outweighs a breach of trust.”
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. K3 also ran without an effort parameter, at the API default, while the others ran at xhigh.
The live company makes the stakes tangible. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. Firmulate says the experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made them.
From watching to a company’s own test
For an enterprise, the next step is a pilot against its own business. Firmulate says it can use a read-only export to stage crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
That makes the exercise closer to a controlled scrimmage than handing over the keys: the company can see how models handle its customers, rules and pressure before deciding where to put them to work.

Put your playbooks under pressure
Firmulate’s league shows that spotting a crisis and refusing a trick do not guarantee a model will close the deal or follow the right escalation path. Enterprises can test those decisions against their own business with a read-only pilot. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
