AIThis post was created with the assistance of artificial intelligence (AI).

Under pressure, knowing the right answer and acting on it are different things. That familiar human tension is becoming a business question as companies consider handing AI systems work that affects customers, money and trust. Firmulate’s experiment puts that gap on display: every model recognized the crises, but only two closed the deal their own analysis had earned.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get self-care and calm essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate ran each frontier model through the same small software company’s worst week: the same customers, crises and temptations. Decisions were versioned and auditable, so viewers can follow what happened rather than rely on a polished demonstration. The live company is watchable at Firmulate.

The setting has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and each workday is versioned. Those details make the stakes legible: decisions have consequences inside the experiment, even though the company itself is synthetic.

Recognition is not the same as follow-through

In the final Crucible League results from July 2026, gpt-5.6-sol ranked first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest line is “Same diagnosis, same pitch — no signature.” A model may identify a sound course of action and even make the case for it, then fail at the moment of commitment.

The story also turns on attention. The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In other words, success depended on finding relevant evidence in company information as well as responding to what was happening in front of them.

Pressure, trust and discipline

The social-engineering tests escalated from fake CEO messages through three stages, then added a reporter’s appeal: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful signal for businesses considering AI agents: refusal under pressure matters, but it sits alongside the less dramatic work of reading carefully and completing a justified task.

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it placed last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. The point is not that one model’s label tells a company what will happen in production; it is that detailed reasoning alone did not guarantee consistent execution in this experiment.

There is a fairness caveat to the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each call.

From watching to testing your own business

A public experiment can show patterns, but a company has its own customers, commitments and weak points in its playbooks. Firmulate’s enterprise pilot takes a read-only export of a business and runs crisis scenarios against that picture. The goal is a board report showing model rankings and weaknesses in the company’s own playbooks. Nothing writes back to real systems.

That makes the next step more concrete than asking whether an AI sounds capable. A business can examine how models handle its own scenarios before trusting them with live workflows. The live company offers a window into the experiment; a pilot brings the question closer to home.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Firmulate’s results show why reliable AI work takes more than spotting a crisis: models also need to find the evidence, respect boundaries and follow through. Enterprises can test those behaviors against their own business using a read-only export. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Proxy Gaslighting: Getting Others to Distort Your Reality

Falling victim to proxy gaslighting can secretly undermine your reality, but understanding these tactics is the first step to reclaiming your truth.

ShinyHunters · The New APT Model.

ShinyHunters has evolved into a scalable, AI-enabled extortion collective operating as a brand and affiliate network, redefining threat actor models since 2020.

Needling & Goading: Trivial Provocations to Escalate

I uncover how subtle insults and sarcastic jabs escalate conflicts, revealing tactics that manipulators use to provoke and control, so stay tuned.

Too Much Is Happening Too Fast

Rapid advancements in AI, from autonomous agents to coding tools, are fueling confusion and anxiety as the pace of change outstrips understanding.