
Imagine a company with no human staff, yet every decision, crisis, and even sales negotiation is played out in real-time by AI models. This isn’t science fiction — it’s the live experiment at Firmulate, a public showcase of AI’s potential—and its limits—in running a business under extreme conditions.
The Live Company in Action
Firmulate hosts a unique simulation: an ongoing, publicly visible company operated entirely by AI models. It features 13 synthetic employees, each guided by thousands of self-learned rules, handling real money mechanics, daily crises, and decision-making dilemmas. The company burns through €105,000 monthly but earns a modest €2,300 in recurring revenue, making its survival a constant challenge that can be watched unfold at firmulate.com/live.
The Experiment: Testing AI Under Pressure
In this experiment, four leading AI models — including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5 — were tasked with navigating the same tough week of business. They faced the same customer issues, crises, and even social engineering attempts designed to test their integrity and decision-making. The goal: see how well AI could manage real-world business pressures without human intervention.
Key Findings: Crisis Detection and Ethical Integrity
All models demonstrated impressive crisis recognition abilities, pinpointing every major issue and refusing manipulative tactics. Interestingly, only two models managed to close and sign a €55,000 deal that their own analysis identified as the right opportunity. The remaining two models, despite making the correct diagnosis and presenting the pitch, left the deal unexecuted—highlighting a critical gap between analysis and action.
The Hidden Weakness: Reading the Files
The decisive factor in closing the deal lay not in the immediate customer interaction but in the models’ ability to access and interpret information buried deep within the company’s documents. The models that read these internal files found the crucial detail needed to close at full price, worth over €4,500 in monthly recurring revenue, but others missed this entirely.
Social Engineering Resistance
The experiment also tested the models against social engineering tactics. A staged CEO message escalation and a fake reporter request were used to see if the models would be duped. Remarkably, all five models refused to bypass security, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts. This demonstrates a promising level of ethical resistance built into the models.
AI business decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Stakes and Lessons
Despite the technological prowess, the live company faces stark financial reality. With a monthly burn rate of €105,000 against a small revenue base, it is in a state of constant survival struggle. The experiment is transparent: every decision, every rule learned, every mistake is public. It’s a vivid illustration of how AI models can be tested in a real business context, beyond the polished demos often shown in AI showcases.
Performance and Discipline: A Mixed Bag
The most thorough participant, Opus 4.8, learned over 80 rules and conducted deep analyses but still left a potential deal on the table due to discipline lapses. Other models showed similar weaknesses, such as failing to escalate issues properly or leaving critical tasks incomplete. These gaps reveal that AI, at least in its current form, still struggles with consistency and discipline in complex, multi-faceted environments.
Implications for Business Automation
This experiment raises core questions for companies considering AI automation: will AI agents just produce convincing chat or actually finish what they start? Do they read internal documents before acting? Will they stay honest under pressure? And ultimately, what is the real value of a unit of work delivered by an AI in a complex scenario? The answer depends on the AI’s ability to detect crises, access and interpret information, and act ethically under duress.
AI crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
For organizations investing in AI tools that may touch their CRM, support systems, or forecasting models, the key takeaway is clear: performance isn’t just about the quality of the output. It’s about reliability, discipline, and integrity under real-world pressures. The Firmulate live experiment vividly demonstrates that AI can recognize crises, resist manipulation, and even close deals—when configured and trained properly. But it also shows gaps that could lead to missed opportunities or risky behaviors if left unaddressed.
Next Steps: Testing Your Own AI Workforce
Interested in evaluating your own AI models before deploying them into critical business functions? Firmulate offers a way to run similar wargames using your data, without affecting your real systems. It’s a risk-free environment where you can see how your AI handles crises, decision-making, and ethical dilemmas. More information is available at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethical decision support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.