
Imagine a skilled negotiator or decision-maker who meticulously studies every document, anticipates every crisis, and refuses to be duped — but still ends up losing the deal. It’s a paradox that unveils a crucial lesson: in the realm of artificial intelligence, sheer effort and volume of rules do not guarantee success. Sometimes, the key isn’t in how much you do, but what you prioritize and how discipline guides your actions. This insight emerges from a groundbreaking public experiment where AI models, tasked with running a simulated company through its worst week, reveal the limits of diligence without strategic focus.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI in Real-World Chaos
At the heart of this story lies a live, observable experiment conducted by Firmulate. Four top-performing AI models were challenged to manage a small software business navigating its most turbulent week. The scenario was real — same customers, same crises, identical temptations to cut corners or manipulate. Every decision made by the models was recorded, versioned, and auditable, creating a transparent window into their decision-making processes.
The Models and Their Scores
- gpt-5.6-sol: scored the highest with a 95 out of 100, demonstrating a remarkable ability to uncover hidden information and close the deal.
- Kimi K3: a newcomer with a score of 93, showed the cleanest discipline and secured the same deal.
- Sonnet 5: scored 88, also closing the deal but with some process slips.
- Fable 5: scored 77, managing to close but with more discipline issues.
Leaders in the league scored well, but all models shared a surprising flaw: despite their diligence, the win was not guaranteed. The critical weakness was in how they handled information deep within the company’s files, not in their immediate crisis recognition or social engineering resistance.
As an affiliate, we earn on qualifying purchases.
Key Findings: The Hidden Weakness of Diligence
All models successfully identified crises and refused manipulative attempts, such as fake CEO messages or staged reporter tricks — even when these escalated over multiple stages. Kimi K3 rationalized its rejection of suspicious requests as a potential impersonation, reflecting a cautious approach.
However, the decisive factor was the ability to recognize and act on buried information. The models that read and understood two document references deep within the company’s files managed to close the deal at full price — worth over €4,583 in monthly recurring revenue.
The Cost of Over-Diligence
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, ultimately finished last in closing the deal. Its discipline slipped under pressure, and it failed to escalate issues appropriately, leaving money on the table. This illustrates that a focus on exhaustive rules and volume without strategic prioritization can hinder actual performance.
business negotiation AI software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Development
This experiment underscores a vital lesson for managers and AI developers alike: diligence and volume of rules do not automatically translate into success. Prioritization, focus, and disciplined escalation are equally — if not more — important. In AI applications touching real business operations, the ability to identify what truly matters and act accordingly can be the difference between closing or losing a deal.
The Bigger Picture: Trust and Impact
Another critical aspect is trust. Despite the models’ technical competence, only two signed the deal they had analyzed and earned. The failure of the others highlights how easily an AI, even one that appears diligent and thorough, can fall short without disciplined focus, especially under pressure. The breach of trust, in this case, was not in the AI’s technical accuracy but in its prioritization and escalation behavior.
As an affiliate, we earn on qualifying purchases.
Learning from the Live Company
The experiment isn’t just theoretical. The live company at firmulate.com runs a sophisticated simulation environment with real money mechanics, 13 synthetic employees, and a public cash countdown — all designed to test AI decision-making in real-time. Observers can watch the AI models in action, making decisions that mirror actual business challenges, and see how they handle crises, temptations, and trust dilemmas.
This setup reveals that the real skill isn’t just in detecting crises but in reading and acting on the most relevant information, avoiding distractions, and maintaining disciplined escalation. It’s a stark reminder that in both artificial and human decision-making, quality over quantity matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI strategic prioritization software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.