Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a company with no human staff, losing €105,000 every month, yet still running a public live experiment that anyone can watch. Welcome to the front lines of AI management, where decision-making, honesty, and resilience are tested in real time. This is not science fiction — it’s the burgeoning world of Firmulate, an open, build-in-public experiment that exposes the realities of AI as a corporate decision-maker.

The Live Experiment: A Company in Crisis, Powered by AI

Firmulate presents a small software business simulated in a live environment, with 13 synthetic employees handling everything from customer crises to strategic decisions. The company is financially strained, burning through €105,000 monthly against a monthly recurring revenue of just €2,300. The cash countdown is public, and every workday, the company’s decision process is versioned, transparent, and auditable.

This setup isn’t just a demo; it’s a real-time testbed for AI models tasked with running a company under extreme stress. Four frontier AI models, including GPT-5.6 and Kimi K3, are pitted against each other, each navigating the same week’s crises, customers, and temptations to cheat or manipulate. The goal: to see if these models can make honest, effective decisions that generate real value.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings from the AI Showdown

All four models successfully identified every crisis and refused every manipulation attempt. This demonstrates a remarkable capacity for situational awareness and integrity under pressure. However, only two of these models managed to close a deal worth €55,000 that their own analysis had earned, effectively balancing ethical decision-making with business outcomes.

The decisive weakness was buried deep in the company’s internal files, not in the immediate customer interactions. When a model read two specific document references, it uncovered a critical piece of information that led to sealing the deal at full price, adding over €4,583 in monthly recurring revenue.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honesty Under Fire: Resisting Social Engineering

The experiment also tested the models against social engineering tactics. Fake CEO messages, staged escalation steps, and even a reporter’s “just one yes/no” request were used to try to bypass the system’s defenses. Impressively, all five models refused every attempt, with Kimi K3 explicitly treating suspicious requests as impersonation risks. This points to an emerging capacity for AI to maintain ethical boundaries even when pressured from multiple angles.

Amazon

AI ethics decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Money, Real Risks: The Firmulate Business

The experiment is anchored in a real-world context: a live, publicly accessible company simulation that burns €105,000 each month while earning only €2,300. The company’s decision-making mechanics are governed by over 680 self-learned rules, all versioned daily. The site, firmulate.com/live.html, offers a window into this ongoing drama, where every decision and crisis is unfold in the open.

This transparency isn’t just for show. The purpose is to evaluate whether AI models can truly replace or augment human judgment, especially in high-stakes, resource-constrained environments. For example, the most thorough participant, Opus 4.8, with over 80 learned rules, showed the deepest analysis but still left close deals on the table and slipped in discipline—highlighting how easily discipline can slip in stressful conditions.

Amazon

AI cybersecurity social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and AI

Many organizations are eager to integrate AI into their operations — from customer support to forecasting. But this experiment shifts focus from AI’s ability to generate convincing chat to its capacity to complete useful, honest work under pressure. Can an AI read your internal files and uncover hidden opportunities? Will it stay honest when facing manipulation? And crucially, what is that work worth?

As the leaderboard shows, even the best models are still imperfect, with scores ranging from 95 for GPT-5.6 to 77 for Fable 5. While GPT-5.6 successfully uncovered the critical internal document and closed the deal, others like Kimi K3 and Sonnet 5 performed almost equally well, with minor slips. Fable 5, however, failed to execute an approved deal, illustrating that rule discipline alone isn’t enough if discipline falters under real stress.

Build in Public, Watch in Real Time

This experiment exemplifies the ethos of open innovation. Anyone can observe the ongoing weekly runs, see how each model handles the same crisis, and even test their own management decisions through the public quiz at firmulate.com/quiz.html. Enterprises can simulate their own business scenarios with a read-only export, helping them evaluate AI before deploying it in the wild.

Ultimately, the core insight is clear: as AI continues to evolve, its ability to handle honest, high-pressure decision-making will be critical. Firmulate’s live, transparent experiment offers a compelling glimpse into the future of AI as a corporate partner — one that must be trusted, resilient, and capable of working amidst chaos and deception.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

In a world where AI influences critical business decisions, transparency and resilience are key. Firmulate’s live experiment shows AI can identify crises, refuse manipulation, and uncover hidden opportunities — but can it truly replace human judgment under pressure? Watch this space for the future of honest, accountable AI decision-making.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Why App-Connected Air Monitors Feel More Useful Over Time

Unlock the lasting benefits of app-connected air monitors as they become more accurate and insightful with use, transforming your indoor environment awareness.

The Smartest Place to Put an Air Purifier in a Small Room

To get the most from your air purifier, place it centrally in…

Why Bedroom Fans Belong in Air-Quality Content Clusters

Great air quality starts with understanding how bedroom fans improve circulation and reduce pollutants—find out why they truly belong in your air-quality content.

How AI Models Show Their True Business Skills in a Crisis — Not in Chat, but in Action

Live AI tests in a real company crisis reveal that only half of top models can finish what they start, underscoring the importance of execution and discipline over chat prowess.