firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Automation leaves the demo room

For readers following AI tools and automation, the most revealing test may not be whether a model can produce a polished answer. It may be whether that model can operate a business when customers are unhappy, money is scarce and an attractive shortcut would violate someone’s trust.

Firmulate turns that question into a public, continuing experiment. Its software company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned.

This is build-in-public pushed toward its logical extreme: the audience is not merely shown product updates or founder reflections. It can watch the company operate while the widening gap between revenue and spending turns survival into a running business story.

AI Automation Mastery: Learn AI Automation, Build Smart Systems, Master No-Code & AI Tools, Boost Productivity, and Turn Your Skills into Income

AI Automation Mastery: Learn AI Automation, Build Smart Systems, Master No-Code & AI Tools, Boost Productivity, and Turn Your Skills into Income

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A terrible week becomes a management test

Firmulate’s Crucible League gave frontier models the same assignment: run the same small software company through its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, making it possible to compare management behavior rather than isolated chat responses.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring: all models identified every crisis and rejected every manipulation attempt. The more instructive result was that only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The winning fact was hiding in plain sight

The deal did not turn on eloquence alone. The decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that followed the references found that fact and won the deal at full price, adding €4,583 in monthly recurring revenue.

That detail should resonate with anyone evaluating automation for real work. Business performance often depends on information scattered across ordinary company records. A model can understand the visible problem and draft a convincing response, yet still fail because it did not investigate the material already available to it. The difference between appearing capable and completing the commercial task was a willingness to read deeper.

Pressure also tested judgment

The models encountered fake messages from a CEO that escalated over three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest description of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance matters because automation is increasingly expected to act around customer records, support conversations and commercial decisions. The experiment suggests that recognizing manipulation may be a strength across the field. It also shows that safety is only part of competent management. Refusing a trap does not compensate for leaving legitimate revenue unsigned.

Thoroughness was not enough

Opus 4.8 provides the sharpest cautionary portrait. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the others, though less strongly.

The contrast exposes a familiar organizational problem. More analysis, more documentation and more accumulated knowledge can coexist with weak execution. The best-looking trail of reasoning is not necessarily the best business outcome if the final commitment is missed or a blocked action is handled poorly.

There is also an important qualification to the ranking: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its 93 result, but it belongs beside the result when readers compare participants.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Artificial Intelligence-Driven Decision Support Framework for Improving Energy Efficiency in Industry and Transportation

Artificial Intelligence-Driven Decision Support Framework for Improving Energy Efficiency in Industry and Transportation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company as a daily accountability feed

Firmulate’s live company makes automation observable over time rather than presenting a single showcase. Visitors can follow the cash countdown, see the accumulating playbook and read what its synthetic employees say. The combination turns abstract questions about AI management into concrete episodes: Was the relevant file read? Was the manipulation refused? Was the earned deal actually signed?

The experiment’s strongest lesson is that spotting a problem is not the same as resolving it. Every model found the crises and resisted the traps, but the field separated when work had to be carried through to a commercial conclusion. For businesses considering AI agents, that gap may be more consequential than fluency. The useful system is not simply the one that sounds like a manager. It is the one that reads carefully, preserves trust, handles obstacles responsibly and finishes the work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Artificial Intelligence for Insurance Fraud Detection: Predictive Models and Risk Analysis

Artificial Intelligence for Insurance Fraud Detection: Predictive Models and Risk Analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Transform Your Team With Grok, The AI That Acts As A Reliable Workmate

xAI announces Grok can now be assigned work as an AI teammate, expanding its role beyond answering prompts, though details remain limited.

Kimi K3: The Gap Closed Six Months Early — And China Stopped Competing On Price

Moonshot’s Kimi K3, a 2.8 trillion-parameter AI, closes the performance gap ahead of schedule and is priced at Western mid-tier levels, challenging assumptions about Chinese AI cost advantages.

Signal: Four Frontier-Class Open Models in Eight Weeks — China’s Release Cadence Is the Story

Chinese labs released four frontier-class open-weight models in just eight weeks, signaling rapid advancement and a shift in AI development cadence.

Radar That Never Blinks: What SAR Actually Does — For Companies, Institutions, And Governments

Explore how Synthetic Aperture Radar (SAR) transforms satellite imaging for companies, institutions, and governments with 24/7, weather-independent capabilities.