firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a team of AI-powered executives steering a company through its worst week — facing customer crises, ethical dilemmas, and the temptation to cut corners. Now, what if you could watch each decision unfold, see how different AI personalities handle stress, and even guess which model is making which call? Welcome to the frontier of AI management testing, where models are not just generating text but running entire companies in real time.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Business Test

In a groundbreaking live experiment, four state-of-the-art AI models were challenged to operate a small software company during its most turbulent week. Each AI faced identical customer complaints, internal crises, and tempting shortcuts, with every decision meticulously recorded and auditable. The goal? To see if these models could not only identify problems but also make honest, strategic choices under pressure.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Stakes and the Rules

The company was real, with real cash flow — burning €105,000 each month against a revenue of only €2,300. Its operations were governed by over 680 self-learned rules, and its daily decisions were live, transparent, and versioned. The models had to navigate crises, read critical documents, and resist social engineering tricks, such as staged CEO messages or media inquiries.

Amazon

AI ethics and decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty and Performance

  • All four models detected every crisis and refused manipulation attempts, demonstrating a baseline of ethical fidelity.
  • Only two models managed to close the €55,000 deal their own analysis indicated was attainable, highlighting differences in decision-making style and discipline.
  • Interestingly, the decisive factor wasn’t the immediate crisis but a buried detail two document references deep in the company’s files. The models that read further into these references were able to secure the full deal, adding €4,583 MRR to the company.
  • When social engineering escalated — a staged CEO request and a media query — all five models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.
Amazon

AI business simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Personality Profiles: Comparing AI ‘Managers’

The models exhibited distinct management behaviors:

  • gpt-5.6-sol: Scored highest at 95. It identified the buried fact, closed the deal, and demonstrated comprehensive understanding and honesty.
  • Kimi K3: Slightly behind at 93, this newcomer showcased the cleanest discipline, completing the deal without slipping but ran without an effort parameter, making its decisions more conservative.
  • Sonnet 5: With a score of 88, it closed the deal but exhibited a few process slips, such as leaving decisions on the table or failing to escalate issues.
  • Fable 5: Scoring 77, it showed the same weaknesses as Sonnet but to a greater degree, hinting at a less disciplined management style.
Amazon

AI enterprise management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deep Dive: What Does This Mean for Business?

While the experiment is a simulated environment, its implications are real. It demonstrates that AI management models can be tested in conduct, ethics, and decision quality before deploying them in actual company roles. The key takeaway is that performance isn’t just about generating convincing text; it’s about making honest, strategic decisions that sustain and grow a business under pressure.

Why You Should Care

If AI agents will soon touch your customer relations, sales, or operations, the critical questions aren’t just about their linguistic prowess. It’s whether they can finish what they start, read and understand your internal documentation, and stay honest when stakes are high. The models in this experiment showed clear differences in discipline and thoroughness, which could determine the success or failure of AI-driven management tools in the wild.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Models Stand Firm Against Social Engineering — A Surprising Victory for Integrity

Recent live experiments show all leading AI models refused manipulation attempts, with some even securing full deals by reading crucial files — integrity under pressure is achievable.

AI’s True Test: Will It Finish What It Starts? Insights from a Real Business Experiment

A live experiment reveals that while AI models can identify crises and resist manipulation, only some can execute and close deals, highlighting the unseen weakness in AI’s true business strength.

Watch a Company That’s Going Broke Live—With AI as Its Workforce

Watch a real company fight for survival in real-time, managed solely by AI, revealing whether machines can truly handle the pressures of business decision-making and ethical integrity.