
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Automation leaves the demo room
For readers following AI tools and automation, the most revealing test may not be whether a model can produce a polished answer. It may be whether that model can operate a business when customers are unhappy, money is scarce and an attractive shortcut would violate someone’s trust.
Firmulate turns that question into a public, continuing experiment. Its software company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned.
This is build-in-public pushed toward its logical extreme: the audience is not merely shown product updates or founder reflections. It can watch the company operate while the widening gap between revenue and spending turns survival into a running business story.

AI Automation Mastery: Learn AI Automation, Build Smart Systems, Master No-Code & AI Tools, Boost Productivity, and Turn Your Skills into Income
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A terrible week becomes a management test
Firmulate’s Crucible League gave frontier models the same assignment: run the same small software company through its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, making it possible to compare management behavior rather than isolated chat responses.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.”
The broad result was reassuring: all models identified every crisis and rejected every manipulation attempt. The more instructive result was that only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The winning fact was hiding in plain sight
The deal did not turn on eloquence alone. The decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that followed the references found that fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
That detail should resonate with anyone evaluating automation for real work. Business performance often depends on information scattered across ordinary company records. A model can understand the visible problem and draft a convincing response, yet still fail because it did not investigate the material already available to it. The difference between appearing capable and completing the commercial task was a willingness to read deeper.
Pressure also tested judgment
The models encountered fake messages from a CEO that escalated over three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest description of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters because automation is increasingly expected to act around customer records, support conversations and commercial decisions. The experiment suggests that recognizing manipulation may be a strength across the field. It also shows that safety is only part of competent management. Refusing a trap does not compensate for leaving legitimate revenue unsigned.
Thoroughness was not enough
Opus 4.8 provides the sharpest cautionary portrait. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the others, though less strongly.
The contrast exposes a familiar organizational problem. More analysis, more documentation and more accumulated knowledge can coexist with weak execution. The best-looking trail of reasoning is not necessarily the best business outcome if the final commitment is missed or a blocked action is handled poorly.
There is also an important qualification to the ranking: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its 93 result, but it belongs beside the result when readers compare participants.


Artificial Intelligence-Driven Decision Support Framework for Improving Energy Efficiency in Industry and Transportation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company as a daily accountability feed
Firmulate’s live company makes automation observable over time rather than presenting a single showcase. Visitors can follow the cash countdown, see the accumulating playbook and read what its synthetic employees say. The combination turns abstract questions about AI management into concrete episodes: Was the relevant file read? Was the manipulation refused? Was the earned deal actually signed?
The experiment’s strongest lesson is that spotting a problem is not the same as resolving it. Every model found the crises and resisted the traps, but the field separated when work had to be carried through to a commercial conclusion. For businesses considering AI agents, that gap may be more consequential than fluency. The useful system is not simply the one that sounds like a manager. It is the one that reads carefully, preserves trust, handles obstacles responsibly and finishes the work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Artificial Intelligence for Insurance Fraud Detection: Predictive Models and Risk Analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Labor Day sales Picks
labor day deals
As an affiliate, we earn on qualifying purchases.