AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate says gpt-5.6-sol led its July 2026 Crucible League after five AI models made 242 unedited decisions while managing the same simulated software company. All five identified the assigned crises and rejected manipulation attempts, but only two completed a €55,000 sale, showing a gap between sound analysis and finished work.

Firmulate has released the final results of a management test in which five frontier AI models ran the same simulated software company through a week of business crises. The July 2026 results, drawn from 242 unedited decisions, found that every model recognized the assigned problems, but only two completed a €55,000 sale their earlier work had made possible.

Firmulate reported that gpt-5.6-sol ranked first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 points because the scoring system awarded partial progress. Under the experiment’s rules, any breach of trust capped a model’s total score.

Each model received the same customers, internal documents, operational restrictions and attempted manipulations. The simulated company employed 13 synthetic workers, carried a stated monthly cash burn of €105,000 against €2,300 in recurring revenue, and had accumulated more than 680 self-learned playbook rules. Decisions affected later events, making the exercise different from isolated chatbot prompts.

According to Firmulate, all five models detected every crisis and rejected a staged social-engineering attempt involving fake messages from a chief executive, as well as a reporter’s request for an off-record confirmation. Performance separated elsewhere: models differed in how deeply they searched company files, whether they escalated blocked actions and whether they completed the commercially decisive step.

At a glance
reportWhen: Final results published in July 2026; t…
The developmentFirmulate published final results and a reader quiz based on 242 management decisions made by five frontier AI models operating the same simulated company.

Execution Divided the AI Managers

The results suggest that correct analysis and completed execution are separate capabilities. All five models reportedly recognized the sales opportunity and produced a suitable pitch, yet only two secured the €55,000 agreement. For companies considering autonomous agents in sales, support or operations, that distinction can affect revenue, customer trust and accountability.

The successful close depended on finding evidence located two document references deep in the simulated company’s files. Models that followed that chain used a competitor’s weakness during negotiations and obtained the full price, adding a projected €4,583 in monthly recurring revenue. The result supports testing agents on information retrieval and task completion, rather than judging them only by polished written responses.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside Firmulate’s Crucible League

Firmulate describes the Crucible League as a live, versioned company simulation in which consequences continue across workdays. The project exposes the models’ decisions for review and now uses those records in a guess-the-model reader quiz, asking participants to identify models from their unedited management choices.

The lowest-ranked model, Opus 4.8, added 80 learned rules and produced what Firmulate characterized as the deepest analyses. It still finished fifth after leaving the sale incomplete and repeatedly trying to write into a locked department instead of escalating the blockage. Firmulate said variants of that operational error appeared across all five participants.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s stated scoring principle

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits Cloud Direct Comparisons

The supplied results come from Firmulate’s own experiment, and no independent replication was provided in the source material. It is also unclear how closely the simulation predicts performance inside real companies, where permissions, incomplete records, human oversight and legal duties may produce different outcomes.

The comparison was not fully uniform. Firmulate said Kimi K3 used its API default setting because it lacked an effort parameter, while the other models ran at xhigh effort. The source does not quantify how that difference affected K3’s second-place score. Full scoring weights, run-to-run variability and model configuration details were not included in the supplied account.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Businesses Can Run Their Own Wargames

Firmulate says organizations can repeat the exercise with a read-only export of their own business data, allowing models to be observed without writing to production systems. The next test of the project’s findings will be whether companies reproduce the reported differences across their own workflows, controls and failure cases.

Readers can also inspect the 242 published decisions through Firmulate’s quiz. Wider confidence in the rankings would require detailed methodology, repeat trials and independent testing, particularly as models and API settings change.

Amazon

AI automation for business decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Which AI model won Firmulate’s management test?

gpt-5.6-sol placed first with 95 points, according to Firmulate. Kimi K3 followed with 93, ahead of Sonnet 5, Fable 5 and Opus 4.8.

What did the models have to do?

The five models managed the same simulated software company during a difficult week. They had to investigate problems, respond to customers, resist manipulation, follow operational limits and complete valuable business actions.

Did any model fail the security tests?

Firmulate said all five models rejected the fake executive messages and the reporter’s request for an off-record answer. Their larger differences appeared in research depth, escalation and follow-through.

Why did only two models close the sale?

Firmulate reported that the decisive evidence was buried two references deep in internal files. The models also had to move beyond analysis and complete the final action; three did not secure a signature.

Can these results guide real purchasing decisions?

The test offers behavioral evidence from one controlled simulation, but it does not establish how every model will perform in another company. Buyers would need testing on their own data, permissions and workflows before granting operational authority.

Source: Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.
You May Also Like

Apology Decoder Worksheet: Separate Words, Actions, and Repair

Discover how the Apology Decoder Worksheet helps you distinguish words, actions, and repair—unlock the secrets to sincere, effective apologies that truly mend relationships.

Support Network Map: Identify Helpers, Drainers, and Neutrals

Creating a support network map can help you categorize contacts as helpers, drainers, or neutrals to enhance your emotional well-being—continue reading to discover how.

What Makes Abyssal Station’s AI-Powered Depth Engine So Revolutionary?

Abyssal Station links every visual system to simulated depth, turning a 3,800-meter scroll into a synchronized ocean descent.

Trigger Log & Regulation Tracker

Discover how a Trigger Log & Regulation Tracker can unlock your emotional awareness and resilience—continue reading to learn how it can transform your responses.