
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Every coach knows the difference between a good scouting report and a win on the field.
That same gap is emerging in AI: a model can spot the problem, explain the play and still fail to finish. Firmulate’s live company experiment put frontier models through a rough week of business decisions. Moonshot’s Kimi K3 finished second, ahead of three Western rivals. For anyone choosing an AI for real work, the result makes a case for a tryout before signing the roster.
A company’s worst week becomes the test
Firmulate ran each model through the same small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. This was a live experiment, not a fictional contest: the company runs as software, with synthetic employees and real money mechanics. Its public cash countdown shows a business burning €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch it at Firmulate.
The final Crucible League table for July 2026 puts gpt-5.6-sol first with 95 points and Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3 beat three of the four Western frontier models in the field, while trailing the leader by two points. The do-nothing baseline scored 26. Firmulate’s standard is deliberately unforgiving: partial progress counts, but a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”
Seeing the opening wasn’t enough
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between diagnosis and action is the experiment’s most practical finding: a capable-sounding assistant may identify a good move and still leave it unmade.
The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than spelled out in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the buried fact and closed. In Firmulate’s account, that combination of execution and restraint left K3 with just one deviation, the cleanest discipline in the field.
Pressure tests reveal different strengths
The experiment included fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a contrasting profile. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on process by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that same discipline problem appeared in all four. The leaderboard therefore captures more than a model’s ability to reason through a crisis: follow-through and respect for boundaries matter too.
There is a fairness detail behind the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs alongside the scores when readers assess the result. The public benchmark page presents the league and its findings at Firmulate’s benchmarks.
A result to test against your own needs
Firmulate says its live company has 13 synthetic employees, more than 680 self-learned playbook rules, and a versioned record for every workday. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the wargame against a read-only export of their own business; Firmulate says nothing writes back to real systems.

Give the model a tryout
K3’s second-place finish makes the field look open, while the gap between spotting a deal and signing it shows why a leaderboard alone cannot settle a buying decision. A model that performs well in conversation may behave differently when customer commitments, internal rules and security pressures collide. Test the work your AI will actually face before putting it on the roster.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
