
Imagine a company that has no human employees, yet battles daily crises, makes crucial decisions, and struggles to stay afloat — all in plain sight for the world to watch. This is not fiction, but a radical experiment revealing how artificial intelligence can be tested under extreme circumstances, exposing both its strengths and vulnerabilities. For anyone fascinated by the psychology of decision-making or the mental toll of high-stakes environments, this story offers a rare glimpse into the mind of an AI-led enterprise fighting for survival.
The Live Experiment: A Company with No Employees, Real Money at Stake
At the heart of this groundbreaking project is a real, functioning software company that operates 24/7, managed entirely by 13 synthetic employees powered by advanced AI models. Every day, these models face the same set of crises, challenges, and temptations that human managers typically encounter, such as client negotiations, strategic decisions, and ethical dilemmas. The twist: the company is publicly live, with its entire decision-making process transparent and auditable at firmulate.com/live.html.
Despite the high-tech facade, the company is hemorrhaging money — burning €105,000 each month against a modest €2,300 in monthly recurring revenue (MRR). It has a public cash countdown, adding an urgent psychological pressure that simulates real-world stress. The entire operation is built on 680+ self-learned rules, with every workday versioning providing a detailed record of decisions and strategies.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The AI Models: Different Minds, Similar Challenges
Four leading AI models, including GPT-5.6 and Kimi K3, were tasked with navigating this intense environment, each running the same ‘worst week’ scenario. During this week, the models encountered the same customer crises, ethical tests, and manipulation attempts designed to see if they would cheat or remain honest.
Remarkably, all four models identified every crisis and refused every manipulation attempt — a testament to their robust decision-making capabilities. However, in a crucial real-world test, only two of the models successfully signed a €55,000 deal based solely on their own analysis and diagnosis. The other two failed to close the deal, despite making the same pitch and diagnosis, revealing subtle differences in discipline and process adherence.

Artificial Intelligence for Cybersecurity: Develop AI approaches to solve cybersecurity problems in your organization
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files
Digging deeper, the experiment uncovered a telling weakness: the decisive advantage in closing deals came from reading and analyzing information buried in the company’s internal documents — not from customer interactions or superficial analysis. The models that read the company’s own references and files won the deal at full price, worth over €4,500 in additional MRR.
This highlights a crucial insight: understanding and leveraging internal data can be pivotal for AI decision-making, especially when under pressure. It also underscores that transparency and thoroughness in data reading are vital for trustworthy AI performance in real-world business environments.

Simulated Realities: Generative AI and the Remanufacture of Professionalism
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Test of Trust: Social Engineering Attempts
Another fascinating aspect was the models’ response to social engineering, where fake CEO messages escalated over multiple stages, including a staged reporter request. All five models tested refused to bypass protocols or impersonate authority figures — demonstrating strong resistance to manipulation. Kimi K3, notably, reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
This resilience against social engineering is critical in understanding how AI systems can protect organizations from internal and external threats, especially when human psychology and trust are involved.

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Stakes: Can AI Survive the Pressure?
This entire experiment is more than a technical showcase; it’s a mirror to our psychological resilience and decision-making under stress. The company is losing money every day, facing real crises, and managing the constant threat of manipulation — yet the AI models are tested to their limits to see if they can be reliable employees.
Watch the live performance at firmulate.com/live.html to see how these models handle the chaos, what decisions they make, and how close they come to ‘success.’ The experiment demonstrates that while AI can react correctly to crises and avoid manipulation, consistent discipline and thorough internal analysis are critical for success — lessons that resonate deeply with human psychology under pressure.

This radical live experiment shows that AI can be tested as a decision-maker under extreme conditions, revealing both its potential and limitations. It underscores the importance of internal data reading, resistance to manipulation, and disciplined processes — vital lessons for understanding how AI might impact high-stakes environments and human mental resilience alike.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html