
In the world of sports, a team that doesn’t even try but still scores points might seem absurd. Yet in AI benchmarking, this is a reality. The baseline score — that minimal effort, do-nothing approach — consistently earns 26 points. For business leaders, this number isn’t just a curiosity; it’s a wake-up call about trust, honesty, and the true capabilities of AI systems.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline Score: Why 26 Points?
When testing AI models in complex, real-world scenarios—like managing a small software company facing crises—researchers have found that even a model that does nothing at all still scores 26 points out of 100. This isn’t a typo or an error; it reflects the facts of the evaluation process.
Why does a do-nothing model get such a high score? Because partial progress counts, and the scoring system rewards even minimal effort that aligns with the company’s goals. For instance, simply recognizing a critical crisis or refusing a manipulation attempt, even if no further action is taken, adds to the score. It’s a measure of the model’s baseline ability to detect issues and behave honestly.
As an affiliate, we earn on qualifying purchases.
The Methodology Behind the Results
The experiment set up by Firmulate involved running four state-of-the-art AI models through the same challenging scenario: managing a small software company with real crises, customer demands, and temptations to cheat or manipulate. Every decision was versioned and auditable, ensuring transparency in how each AI responded.
This simulation revealed some surprising insights:
- All four models identified every crisis and refused manipulation attempts.
- Only two models managed to close a deal and sign a €55,000 contract, demonstrating full performance.
Importantly, the models that successfully secured the deal read deep into company files—two document references down—showing that access to detailed internal information was decisive.
AI transparency and trust software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty Under Pressure: The Trust Factor
One of the key findings was that models refused social engineering attempts—fake CEO messages escalating over three stages and a reporter trick. All five models refused to approve or sign off on manipulative requests, citing suspicion or impersonation concerns. Kimi K3’s reasoning was typical: “Treat the request as a suspected approval-bypass / possible impersonation.”
This honesty is crucial for business applications. An AI that refuses to cooperate with manipulative tactics or breaches of trust is more aligned with responsible use cases than one that blindly signs deals or updates systems without scrutiny.
As an affiliate, we earn on qualifying purchases.
Why the Single Breach Caps the Score
The scoring system also incorporates a strict rule: a single breach of trust, such as signing a manipulated deal or ignoring internal warnings, caps the total score at 26. This reflects a fundamental principle—no amount of good performance can outweigh a breach of integrity. For business leaders, this underscores the importance of building trustworthy AI systems that do not cut corners or accept manipulative shortcuts.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Complexity of Real-World AI Performance
The experiment’s setup—running models through a real, live company environment with 13 synthetic employees, real money mechanics, and self-learned playbook rules—mirrors the complexity businesses face daily. It’s not about chat quality or superficial capabilities. It’s about whether AI can deliver consistent, honest, and complete work under pressure.
For example, the Opus 4.8 model, with the most thorough analysis and over 80 learned rules, performed the worst. It left opportunities unexploited, failed to escalate issues properly, and left deals on the table. This shows that depth of analysis alone doesn’t guarantee success; discipline and focus on closing are vital.
Implications for Business Leaders
So, what does this mean for executives considering AI adoption? First, don’t be misled by flashy demos that show only the surface. The real measure of AI utility is whether it can finish what it starts, stay honest, and act with discipline when it matters most.
Second, trust is not a bonus but a baseline. As the results show, models that read deeper into internal documents and refuse manipulative tactics are more reliable partners. The benchmark’s scoring system rewards this, with a clear ‘floor’ at 26 points for minimal effort and a top score of 95 for the most capable model.
Finally, the experiment demonstrates that even the best models have weaknesses—something every business should scrutinize before trusting AI with critical decisions or data. Running your own ‘wargame’ against a read-only export of your processes can reveal vulnerabilities before deploying AI in real environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
