
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The test of an AI manager is what it does under pressure
In psychology, recognizing a problem and acting on it are different things. Firmulate’s company experiment puts that gap on display: every model spotted every crisis and rejected every manipulation attempt, but only two signed the deal their own analysis had earned. A polished diagnosis, it turns out, does not guarantee follow-through.
A newcomer enters an open league
Firmulate ran frontier AI models through the same worst week at a small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. In the final July 2026 Crucible league, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; the stated rule is that partial progress counts, but a single breach of trust caps the total.
K3’s performance was practical as well as disciplined: it found the buried security weakness, won the €55,000 deal worth +€4,583 in monthly recurring revenue, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field. The fairness caveat matters: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.
Finding the fact—and finishing the job
The decisive competitor weakness was hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price. That makes the result a test of attention and follow-through as much as crisis recognition: the useful clue was available, but only if a model looked beyond the immediate prompt.
Across the field, all models identified every crisis and refused every manipulation attempt. Yet only two signed the deal their analysis had justified. Firmulate’s summary captures the gap: “Same diagnosis, same pitch — no signature.” K3 was one of the models that closed. Its result suggests why judging an AI workforce from fluent answers alone can miss the harder question: whether it completes consequential work when a decision is required.
Resisting pressure, keeping discipline
The social engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background”. All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear response to a pressure tactic designed to make a sensitive request seem routine.
Strong results did not mean flawless execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four. The contrast is a reminder that extensive analysis and procedural care do not automatically produce a sound close.
A watchable company, not a chat demo
The live Firmulate company has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The experiment is watchable at Firmulate. Its premise is to measure management quality, not chat quality, by letting models operate through a company’s ordinary pressures.
A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; the pilot writes nothing back to real systems. That offers a way to examine model behavior against familiar workflows before giving an AI agent access to them.

Test the behavior you will depend on
Kimi K3’s second-place finish makes the league look open: it beat three of the four Western frontier models in this experiment, while gpt-5.6-sol took first. The result does not settle which model will suit every company. It does show why choosing on reputation or a chat demonstration alone is a bet. Look at whether a model finds critical information, closes the work it has earned and holds its boundaries under pressure. See the full benchmark findings before making that call.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
