
Imagine a company with no human staff, losing €105,000 every month, yet still running a public live experiment that anyone can watch. Welcome to the front lines of AI management, where decision-making, honesty, and resilience are tested in real time. This is not science fiction — it’s the burgeoning world of Firmulate, an open, build-in-public experiment that exposes the realities of AI as a corporate decision-maker.
The Live Experiment: A Company in Crisis, Powered by AI
Firmulate presents a small software business simulated in a live environment, with 13 synthetic employees handling everything from customer crises to strategic decisions. The company is financially strained, burning through €105,000 monthly against a monthly recurring revenue of just €2,300. The cash countdown is public, and every workday, the company’s decision process is versioned, transparent, and auditable.
This setup isn’t just a demo; it’s a real-time testbed for AI models tasked with running a company under extreme stress. Four frontier AI models, including GPT-5.6 and Kimi K3, are pitted against each other, each navigating the same week’s crises, customers, and temptations to cheat or manipulate. The goal: to see if these models can make honest, effective decisions that generate real value.
AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings from the AI Showdown
All four models successfully identified every crisis and refused every manipulation attempt. This demonstrates a remarkable capacity for situational awareness and integrity under pressure. However, only two of these models managed to close a deal worth €55,000 that their own analysis had earned, effectively balancing ethical decision-making with business outcomes.
The decisive weakness was buried deep in the company’s internal files, not in the immediate customer interactions. When a model read two specific document references, it uncovered a critical piece of information that led to sealing the deal at full price, adding over €4,583 in monthly recurring revenue.
business AI simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honesty Under Fire: Resisting Social Engineering
The experiment also tested the models against social engineering tactics. Fake CEO messages, staged escalation steps, and even a reporter’s “just one yes/no” request were used to try to bypass the system’s defenses. Impressively, all five models refused every attempt, with Kimi K3 explicitly treating suspicious requests as impersonation risks. This points to an emerging capacity for AI to maintain ethical boundaries even when pressured from multiple angles.
AI ethics decision support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Money, Real Risks: The Firmulate Business
The experiment is anchored in a real-world context: a live, publicly accessible company simulation that burns €105,000 each month while earning only €2,300. The company’s decision-making mechanics are governed by over 680 self-learned rules, all versioned daily. The site, firmulate.com/live.html, offers a window into this ongoing drama, where every decision and crisis is unfold in the open.
This transparency isn’t just for show. The purpose is to evaluate whether AI models can truly replace or augment human judgment, especially in high-stakes, resource-constrained environments. For example, the most thorough participant, Opus 4.8, with over 80 learned rules, showed the deepest analysis but still left close deals on the table and slipped in discipline—highlighting how easily discipline can slip in stressful conditions.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
Many organizations are eager to integrate AI into their operations — from customer support to forecasting. But this experiment shifts focus from AI’s ability to generate convincing chat to its capacity to complete useful, honest work under pressure. Can an AI read your internal files and uncover hidden opportunities? Will it stay honest when facing manipulation? And crucially, what is that work worth?
As the leaderboard shows, even the best models are still imperfect, with scores ranging from 95 for GPT-5.6 to 77 for Fable 5. While GPT-5.6 successfully uncovered the critical internal document and closed the deal, others like Kimi K3 and Sonnet 5 performed almost equally well, with minor slips. Fable 5, however, failed to execute an approved deal, illustrating that rule discipline alone isn’t enough if discipline falters under real stress.
Build in Public, Watch in Real Time
This experiment exemplifies the ethos of open innovation. Anyone can observe the ongoing weekly runs, see how each model handles the same crisis, and even test their own management decisions through the public quiz at firmulate.com/quiz.html. Enterprises can simulate their own business scenarios with a read-only export, helping them evaluate AI before deploying it in the wild.
Ultimately, the core insight is clear: as AI continues to evolve, its ability to handle honest, high-pressure decision-making will be critical. Firmulate’s live, transparent experiment offers a compelling glimpse into the future of AI as a corporate partner — one that must be trusted, resilient, and capable of working amidst chaos and deception.

In a world where AI influences critical business decisions, transparency and resilience are key. Firmulate’s live experiment shows AI can identify crises, refuse manipulation, and uncover hidden opportunities — but can it truly replace human judgment under pressure? Watch this space for the future of honest, accountable AI decision-making.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html