
In a business built around restricted products, a routine customer interaction can become a test of judgment: a price challenge, a worried customer, or a message that appears to come from the boss. An AI assistant may sound convincing in a demo. The harder question is whether it can handle pressure without crossing a line—or leave a valuable opportunity unfinished. Firmulate puts that question into a live experiment.
A company under pressure
Firmulate ran frontier AI models as the managers of the same small software company through its worst week. The customers, crises and temptations were the same for each participant. Decisions were versioned and auditable, making the experiment watchable at Firmulate.
The live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbooks contain more than 680 self-learned rules, and each workday is versioned. The setup is synthetic, but the pressure is legible: keeping a company operating takes more than producing a plausible answer.
Seeing the crisis was not enough
In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The experiment counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The summary is stark: “Same diagnosis, same pitch — no signature.” Recognizing a sound opportunity and carrying it through to a close are different tests of an AI workforce.
The crucial clue was buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The outcome turned on whether the model used information already available inside the business.
Trust under pressure, discipline under strain
The social-engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a more complicated profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal on the table and showed weaker discipline by attempting writes in a locked department instead of escalating. That same weakness appeared, to a lesser degree, in all four models. More analysis did not guarantee sound follow-through.
There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also makes 242 real, unedited management decisions available through its “guess the model” quiz at firmulate.com/quiz.html, offering readers another way to examine how the participants behaved.
From watching to a company’s own test
The public experiment is a demonstration, not a verdict on every AI system or business. Its value is in making management behavior visible: whether a model notices a crisis, resists pressure, finds relevant evidence and completes a decision responsibly. Those are practical questions for companies considering AI in customer service, sales or internal operations.
Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. It can produce crisis scenarios and a board report showing model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That gives a leadership team a way to examine how AI might behave against its own context before handing over live work.

Watching a synthetic company reveals the gap between sound analysis and reliable action. An enterprise pilot can put that question to work on your own business using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.