AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The test of an AI manager is what it does under pressure

In psychology, recognizing a problem and acting on it are different things. Firmulate’s company experiment puts that gap on display: every model spotted every crisis and rejected every manipulation attempt, but only two signed the deal their own analysis had earned. A polished diagnosis, it turns out, does not guarantee follow-through.

A newcomer enters an open league

Firmulate ran frontier AI models through the same worst week at a small software company, with the same customers, crises and temptations. Decisions were versioned and auditable. In the final July 2026 Crucible league, Moonshot’s Kimi K3 scored 93, placing second behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; the stated rule is that partial progress counts, but a single breach of trust caps the total.

K3’s performance was practical as well as disciplined: it found the buried security weakness, won the €55,000 deal worth +€4,583 in monthly recurring revenue, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field. The fairness caveat matters: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Finding the fact—and finishing the job

The decisive competitor weakness was hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price. That makes the result a test of attention and follow-through as much as crisis recognition: the useful clue was available, but only if a model looked beyond the immediate prompt.

Across the field, all models identified every crisis and refused every manipulation attempt. Yet only two signed the deal their analysis had justified. Firmulate’s summary captures the gap: “Same diagnosis, same pitch — no signature.” K3 was one of the models that closed. Its result suggests why judging an AI workforce from fluent answers alone can miss the harder question: whether it completes consequential work when a decision is required.

Resisting pressure, keeping discipline

The social engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background”. All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear response to a pressure tactic designed to make a sensitive request seem routine.

Strong results did not mean flawless execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four. The contrast is a reminder that extensive analysis and procedural care do not automatically produce a sound close.

A watchable company, not a chat demo

The live Firmulate company has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The experiment is watchable at Firmulate. Its premise is to measure management quality, not chat quality, by letting models operate through a company’s ordinary pressures.

A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; the pilot writes nothing back to real systems. That offers a way to examine model behavior against familiar workflows before giving an AI agent access to them.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the behavior you will depend on

Kimi K3’s second-place finish makes the league look open: it beat three of the four Western frontier models in this experiment, while gpt-5.6-sol took first. The result does not settle which model will suit every company. It does show why choosing on reputation or a chat demonstration alone is a bet. Look at whether a model finds critical information, closes the work it has earned and holds its boundaries under pressure. See the full benchmark findings before making that call.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Projection: When Accusations Reveal the Accuser

Just when you think accusations reveal truths about others, they might instead expose your own insecurities—discover the deeper implications of projection.

Death of an Honor Code

Princeton University has announced the return of proctored exams amid rising AI-facilitated cheating, marking the end of its longstanding honor system.

Revealing The Hidden Majority: How 1% Of Social Media Users Shape Our Perception Of Lifestyle Trends

Research reveals that roughly 1% of social media users generate most online content, shaping perceptions while 90% remain silent. This impacts how we view online opinions.