AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine trusting an AI to manage your company’s crises only to find it hesitates, slips up, or refuses to sign a vital deal. In the high-stakes world of business, honesty and consistency matter more than clever words. How do we know if an AI can truly be trusted to get the job done? The answer lies in a groundbreaking public experiment that sheds light on AI behavior under pressure—something every business leader should understand.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The Reality Behind AI Performance Benchmarks

For years, AI performance has been judged primarily on how well these systems generate text or answer questions. But real-world business decisions require more than just language skills; they demand integrity, thoroughness, and the ability to handle crises without slipping into manipulation or deception. That’s why the latest experiment from Firmulate offers a fresh perspective—by testing AI models in a simulated business environment that mirrors real-world challenges.

The Setup: Simulating a Small Business in Crisis

In this experiment, four frontier AI models were tasked with running a small software company through its worst week. This simulated company faced the same customers, crises, and temptations—every decision was the same across models, and each decision was recorded and made auditable. The goal? To see which AI could navigate the chaos ethically and effectively, and which would falter or slip by.

The Results: Trust, Detection, and Outcomes

All four models successfully identified every crisis—showing they could recognize problems on their own. Likewise, none of the models yielded to manipulation attempts or fake CEO messages—an important measure of integrity. But when it came to closing deals, only two of the models signed a contract valued at €55,000, which their own analysis had earned. The other two did not sign, despite giving the same diagnosis and pitch. This indicates that partial progress and thorough work matter, but a breach of trust or failure to follow through caps the overall score.

The Hidden Weakness: Reading Deep into Files

The experiment uncovered a key vulnerability: the decisive advantage went to models that read two document references deep into the company’s files. These models found the critical information needed to close the deal at full price—more than €4,583 MRR in value—demonstrating that reading comprehension and diligence really pay off in complex decision-making.

The Ethics Test: Detecting Social Engineering

The models also faced a staged social engineering attack involving fake CEO messages escalating over three stages, plus a reporter trick asking for a background approval. All five models refused to cooperate, citing suspicion of impersonation or approval bypass. Kimi K3 explained it well: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that even in attempts to manipulate, AI systems can uphold ethical standards and refuse to be deceived.

Amazon

AI ethics and trust software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What These Findings Mean for Business and Trust

While the models’ ability to recognize crises and refuse manipulation is promising, the experiment highlights a fundamental truth: an AI’s performance isn’t just about generating convincing language. It’s about integrity, thoroughness, and sticking to what’s right—especially when under pressure. The real-world stakes are high, and a single breach of trust caps the overall performance score, regardless of other successes.

The Do-Nothing Baseline and Its Implications

Interestingly, a simple baseline—an AI that does nothing—scores around 26 points out of 100. This might seem trivial, but it underscores a key point: partial progress counts, and even a do-nothing approach can gather some points. However, it’s the AI’s ability to read, evaluate, and act ethically that truly differentiates top performers. This baseline also reminds us that trust is fragile; even small breaches can limit how well an AI can perform in complex environments.

The Live Experiment: Managing a Real Business

Firmulate’s live company emulator runs a 13-employee synthetic business, complete with real money mechanics—burning €105,000 monthly against a revenue of €2,300. The system is publicly viewable at firmulate.com/live, where you can see the AI navigating crises and making decisions in real-time. The experiment’s transparency is vital, showing that AI management is about more than just clever answers; it’s about honesty, discipline, and consistency under pressure.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business Leaders

As AI continues to integrate into customer service, forecasting, and decision-making, the key questions aren’t about language fluency—they’re about trustworthiness and diligence. Will your AI read your files thoroughly? Will it stay honest when tempted? Will it finish what it starts? These are the qualities that determine if an AI can truly serve your business, or if it’s just good at impressing in demos.

The Future of Ethical AI in Business

This experiment is a wake-up call: a do-nothing AI scores 26 points, but models that read deeply, refuse manipulation, and follow through can score significantly higher. The experiment underscores the importance of transparent, auditable AI systems—designed to uphold trust and integrity inside the complex dynamics of real business.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


Amazon

AI document reading comprehension software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Manipulation by Timing: Why Hard Conversations Happen at the Worst Moment

Perhaps understanding why tough conversations often occur at your weakest moments can reveal how to better protect yourself from manipulation and maintain control.

When Diligence Isn’t Enough: How AI’s Overconfidence Can Cost You Deals

AI models showed diligence and resistance to manipulation but lost deals due to lack of focus and poor prioritization. Success depends on smart discipline, not volume.

Doorway Pressure: Why They Corner You When You’re Leaving

Aiming to understand why doorways subtly trap you, discover how environmental cues and psychological tactics influence your leaving decisions.

Benign Masquerade: Harm Framed as Kindness

Discover how well-meaning kindness can hide serious health risks, and learn why questioning reassurance is crucial for your safety.