
Imagine a team of AI-powered executives steering a company through its worst week — facing customer crises, ethical dilemmas, and the temptation to cut corners. Now, what if you could watch each decision unfold, see how different AI personalities handle stress, and even guess which model is making which call? Welcome to the frontier of AI management testing, where models are not just generating text but running entire companies in real time.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Business Test
In a groundbreaking live experiment, four state-of-the-art AI models were challenged to operate a small software company during its most turbulent week. Each AI faced identical customer complaints, internal crises, and tempting shortcuts, with every decision meticulously recorded and auditable. The goal? To see if these models could not only identify problems but also make honest, strategic choices under pressure.
AI decision-making management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Stakes and the Rules
The company was real, with real cash flow — burning €105,000 each month against a revenue of only €2,300. Its operations were governed by over 680 self-learned rules, and its daily decisions were live, transparent, and versioned. The models had to navigate crises, read critical documents, and resist social engineering tricks, such as staged CEO messages or media inquiries.
AI ethics and decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Honesty and Performance
- All four models detected every crisis and refused manipulation attempts, demonstrating a baseline of ethical fidelity.
- Only two models managed to close the €55,000 deal their own analysis indicated was attainable, highlighting differences in decision-making style and discipline.
- Interestingly, the decisive factor wasn’t the immediate crisis but a buried detail two document references deep in the company’s files. The models that read further into these references were able to secure the full deal, adding €4,583 MRR to the company.
- When social engineering escalated — a staged CEO request and a media query — all five models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.
As an affiliate, we earn on qualifying purchases.
The Personality Profiles: Comparing AI ‘Managers’
The models exhibited distinct management behaviors:
- gpt-5.6-sol: Scored highest at 95. It identified the buried fact, closed the deal, and demonstrated comprehensive understanding and honesty.
- Kimi K3: Slightly behind at 93, this newcomer showcased the cleanest discipline, completing the deal without slipping but ran without an effort parameter, making its decisions more conservative.
- Sonnet 5: With a score of 88, it closed the deal but exhibited a few process slips, such as leaving decisions on the table or failing to escalate issues.
- Fable 5: Scoring 77, it showed the same weaknesses as Sonnet but to a greater degree, hinting at a less disciplined management style.
AI enterprise management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deep Dive: What Does This Mean for Business?
While the experiment is a simulated environment, its implications are real. It demonstrates that AI management models can be tested in conduct, ethics, and decision quality before deploying them in actual company roles. The key takeaway is that performance isn’t just about generating convincing text; it’s about making honest, strategic decisions that sustain and grow a business under pressure.
Why You Should Care
If AI agents will soon touch your customer relations, sales, or operations, the critical questions aren’t just about their linguistic prowess. It’s whether they can finish what they start, read and understand your internal documentation, and stay honest when stakes are high. The models in this experiment showed clear differences in discipline and thoroughness, which could determine the success or failure of AI-driven management tools in the wild.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.