firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every coach knows the difference between a good scouting report and a win on the field.

That same gap is emerging in AI: a model can spot the problem, explain the play and still fail to finish. Firmulate’s live company experiment put frontier models through a rough week of business decisions. Moonshot’s Kimi K3 finished second, ahead of three Western rivals. For anyone choosing an AI for real work, the result makes a case for a tryout before signing the roster.

A company’s worst week becomes the test

Firmulate ran each model through the same small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. This was a live experiment, not a fictional contest: the company runs as software, with synthetic employees and real money mechanics. Its public cash countdown shows a business burning €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch it at Firmulate.

The final Crucible League table for July 2026 puts gpt-5.6-sol first with 95 points and Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3 beat three of the four Western frontier models in the field, while trailing the leader by two points. The do-nothing baseline scored 26. Firmulate’s standard is deliberately unforgiving: partial progress counts, but a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”

Seeing the opening wasn’t enough

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between diagnosis and action is the experiment’s most practical finding: a capable-sounding assistant may identify a good move and still leave it unmade.

The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than spelled out in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the buried fact and closed. In Firmulate’s account, that combination of execution and restraint left K3 with just one deviation, the cleanest discipline in the field.

Pressure tests reveal different strengths

The experiment included fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a contrasting profile. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on process by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that same discipline problem appeared in all four. The leaderboard therefore captures more than a model’s ability to reason through a crisis: follow-through and respect for boundaries matter too.

There is a fairness detail behind the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs alongside the scores when readers assess the result. The public benchmark page presents the league and its findings at Firmulate’s benchmarks.

A result to test against your own needs

Firmulate says its live company has 13 synthetic employees, more than 680 self-learned playbook rules, and a versioned record for every workday. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the wargame against a read-only export of their own business; Firmulate says nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Give the model a tryout

K3’s second-place finish makes the field look open, while the gap between spotting a deal and signing it shows why a leaderboard alone cannot settle a buying decision. A model that performs well in conversation may behave differently when customer commitments, internal rules and security pressures collide. Test the work your AI will actually face before putting it on the roster.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Models Run a Business Under Pressure? Watch Them Decide in Real Time

Watch AI models run a real company through its toughest week, revealing their decision-making styles, honesty, and discipline — crucial factors for future AI management tools.

Exact Margin: Giants By 25+

Betting markets indicate a strong consensus that the Giants will win their next game by at least 25 points, with a current market probability of 94%.

Tour Italy, Sicily In 2027 With Barnard Griffin

Barnard Griffin announces a guided tour of Italy and Sicily scheduled for 2027, offering travelers a curated experience of these historic regions.

AI’s True Test: Will It Finish What It Starts? Insights from a Real Business Experiment

A live experiment reveals that while AI models can identify crises and resist manipulation, only some can execute and close deals, highlighting the unseen weakness in AI’s true business strength.