AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

When a company faces its worst week—crises piling up, temptations to cheat, and pressure mounting—how does leadership respond? Human psychology suggests that stress can push even the most honest to cut corners or abandon their principles. But what if artificial intelligence could serve as a mirror—not just of decision-making, but of moral resilience? A groundbreaking experiment with AI models running a real company has revealed surprising insights into integrity, discipline, and the limits of automation under pressure.

The Setup: Simulating a Company’s Darkest Hour

In a live experiment, four advanced AI models were tasked with managing the same small software company during its worst week. The scenario included genuine crises—customer issues, financial temptations, and manipulative social engineering attacks—every challenge a real business might face. Every decision was recorded, versioned, and auditable, ensuring transparency and accountability in how each AI responded.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Findings: Honesty and Competence Under Stress

Remarkably, all four models identified every crisis and refused every attempt to manipulate them. This shows that AI, at the very least, can be trained to recognize threats and resist unethical pressure—an echo of the human ideal of moral fortitude. However, when it came to closing a key deal worth €55,000, only two models succeeded in executing their own analysis and signing the agreement. The other two detected the opportunity but failed to act on it, leaving potential revenue on the table despite having identified the same facts.

Amazon

AI cybersecurity social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

The critical difference lay in the models’ ability to interpret the company’s own internal documents. The winning models read two document references deep into the company’s files—information that was buried but vital. Those who accessed and understood this internal data closed the deal at full price, adding €4,583 Monthly Recurring Revenue (MRR). This demonstrates that true operational competence and moral resolve are linked to thoroughness and attention to detail—qualities that AI can measure but are often overlooked in superficial assessments like chat demos.

Amazon

enterprise AI compliance and transparency tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Test of Integrity: Social Engineering and Pressure

Social engineering attacks—fake CEO messages and a reporter’s subtle request—were staged to test whether AI could be manipulated. All five models refused to engage. Kimi K3 explicitly treated such requests as suspicious, citing the risk of impersonation. This resilience under pressure is crucial for real-world AI deployments, where unethical manipulations are commonplace and can cause severe harm if not appropriately thwarted.

Amazon

AI ethics and resilience training programs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business: A Live, Losing Company

The experiment’s backdrop is a real company—13 synthetic employees managing real money mechanics, burning €105,000 per month against a revenue stream of only €2,300. Every day, the company evolves with over 680 self-learned rules, and its decisions are transparent for observers. This provides a tangible context for understanding how AI decision-making translates into actual business outcomes, not just theoretical chat performance.

The Lessons: Disciplined AI Is a Trustworthy Partner

Interestingly, the most disciplined model, Opus 4.8, with over 80 learned rules and deep analysis, failed to close the deal—leaving the opportunity unexecuted due to process slips. This highlights that being thorough and rule-abiding is not enough; decisive action and discipline in execution matter just as much. The same weakness appeared across all models, hinting at a fundamental challenge: AI can recognize problems and resist pressure but may still falter in completing critical tasks without targeted design.

What This Means for Business and Psychology

For those attuned to the intricacies of mental resilience and integrity, this experiment underscores a vital point: the true test of a system—human or machine—is not how well it recognizes problems in a conversation but whether it follows through under real-world pressures. AI models that can read internal documents, resist manipulation, and execute decisions reliably demonstrate a form of moral consistency that is often elusive in human behavior, especially under stress.

The Bigger Picture: Beyond Chat Skills

In today’s AI landscape, much emphasis is placed on chat demos and superficial performance metrics. But this experiment shows that the real measure of an AI’s usefulness is its capacity for discipline, thoroughness, and integrity—traits that are critical when actual money, reputation, and trust are at stake.

Explore the Future: Run Your Own Business Wargame

Interested in testing your organization’s resilience? Firmulate offers a platform where companies can simulate their own worst weeks with AI. Using a read-only export of your business, you can see how AI models handle crises, temptations, and moral dilemmas—without risking real operations. Visit firmulate.com to learn how to run your own digital twin and discover whether your AI workforce can truly be trusted to deliver results when it matters most.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Too Much Is Happening Too Fast

Rapid advancements in AI, from autonomous agents to coding tools, are fueling confusion and anxiety as the pace of change outstrips understanding.

The cleaner cap table. Why Anthropic’s public-benefit structure dodges OpenAI’s charitable-trust problem — and trades it for a governance question of its own.

Anthropic’s public-benefit corporate structure provides a simpler ownership model compared to OpenAI’s charitable trust, raising governance questions.

Evidence Dumps”: Cherry‑Picked Proof to Win the Frame

Could evidence dumps be manipulated to unfairly sway justice, leaving you wondering how to uncover the full truth?

Triangulation: How Third Parties Get Used to Control

Sow the seeds of understanding about triangulation and discover how to reclaim your relationships from third-party manipulation. The truth might surprise you.