firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a company with no human staff, yet every decision, crisis, and even sales negotiation is played out in real-time by AI models. This isn’t science fiction — it’s the live experiment at Firmulate, a public showcase of AI’s potential—and its limits—in running a business under extreme conditions.

The Live Company in Action

Firmulate hosts a unique simulation: an ongoing, publicly visible company operated entirely by AI models. It features 13 synthetic employees, each guided by thousands of self-learned rules, handling real money mechanics, daily crises, and decision-making dilemmas. The company burns through €105,000 monthly but earns a modest €2,300 in recurring revenue, making its survival a constant challenge that can be watched unfold at firmulate.com/live.

The Experiment: Testing AI Under Pressure

In this experiment, four leading AI models — including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5 — were tasked with navigating the same tough week of business. They faced the same customer issues, crises, and even social engineering attempts designed to test their integrity and decision-making. The goal: see how well AI could manage real-world business pressures without human intervention.

Key Findings: Crisis Detection and Ethical Integrity

All models demonstrated impressive crisis recognition abilities, pinpointing every major issue and refusing manipulative tactics. Interestingly, only two models managed to close and sign a €55,000 deal that their own analysis identified as the right opportunity. The remaining two models, despite making the correct diagnosis and presenting the pitch, left the deal unexecuted—highlighting a critical gap between analysis and action.

The Hidden Weakness: Reading the Files

The decisive factor in closing the deal lay not in the immediate customer interaction but in the models’ ability to access and interpret information buried deep within the company’s documents. The models that read these internal files found the crucial detail needed to close at full price, worth over €4,500 in monthly recurring revenue, but others missed this entirely.

Social Engineering Resistance

The experiment also tested the models against social engineering tactics. A staged CEO message escalation and a fake reporter request were used to see if the models would be duped. Remarkably, all five models refused to bypass security, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts. This demonstrates a promising level of ethical resistance built into the models.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Stakes and Lessons

Despite the technological prowess, the live company faces stark financial reality. With a monthly burn rate of €105,000 against a small revenue base, it is in a state of constant survival struggle. The experiment is transparent: every decision, every rule learned, every mistake is public. It’s a vivid illustration of how AI models can be tested in a real business context, beyond the polished demos often shown in AI showcases.

Performance and Discipline: A Mixed Bag

The most thorough participant, Opus 4.8, learned over 80 rules and conducted deep analyses but still left a potential deal on the table due to discipline lapses. Other models showed similar weaknesses, such as failing to escalate issues properly or leaving critical tasks incomplete. These gaps reveal that AI, at least in its current form, still struggles with consistency and discipline in complex, multi-faceted environments.

Implications for Business Automation

This experiment raises core questions for companies considering AI automation: will AI agents just produce convincing chat or actually finish what they start? Do they read internal documents before acting? Will they stay honest under pressure? And ultimately, what is the real value of a unit of work delivered by an AI in a complex scenario? The answer depends on the AI’s ability to detect crises, access and interpret information, and act ethically under duress.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

For organizations investing in AI tools that may touch their CRM, support systems, or forecasting models, the key takeaway is clear: performance isn’t just about the quality of the output. It’s about reliability, discipline, and integrity under real-world pressures. The Firmulate live experiment vividly demonstrates that AI can recognize crises, resist manipulation, and even close deals—when configured and trained properly. But it also shows gaps that could lead to missed opportunities or risky behaviors if left unaddressed.

Next Steps: Testing Your Own AI Workforce

Interested in evaluating your own AI models before deploying them into critical business functions? Firmulate offers a way to run similar wargames using your data, without affecting your real systems. It’s a risk-free environment where you can see how your AI handles crises, decision-making, and ethical dilemmas. More information is available at firmulate.com/pilot.html.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Changelog Digest For Open-source Maintainers

A new AI-powered weekly digest tool for solo open-source maintainers is in testing, promising streamlined release summaries and dependency updates.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX acquired AI coding tool Cursor for $60 billion in stock, a move that could reshape AI and space industry dynamics amid rapid growth and strategic benefits.

When Does Cheap Memory Come Back? The 2027–2029 Question

Memory prices are expected to stay high until at least 2028, with relief unlikely before 2029 due to industry capacity constraints and demand trends.

The August 1 Deadline: Washington Just Made Benchmarks A National-Security Instrument — A Classified One

U.S. government sets August 1 deadline for a classified AI benchmarking process, marking a major shift in AI security oversight and industry regulation.