
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Leaderboard Says Your Agent Is Brilliant. Your P&L Might Disagree.
If you follow AI tooling, you know the ritual by now: a new model drops, it rockets up a coding benchmark, someone posts a screenshot of an impossibly elegant function, and the procurement conversation starts. But the benchmarks that decide which agents get hired all measure the same thing — how well a model answers. They don’t measure what happens after the answer: whether the agent finishes what it starts, whether it reads your files before acting, and whether it stays honest when the week goes sideways.
That gap now has a number attached to it, thanks to a live experiment run by Firmulate, which runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality.
As an affiliate, we earn on qualifying purchases.
Four Models, One Terrible Week
The setup is brutally simple. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision was versioned and auditable, so nothing about the outcome is vibes.
The final league table from the July 2026 run:
- gpt-5.6-sol — 95 points. Found the buried fact, closed the deal, described as “the complete performance.”
- Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
- Sonnet 5 — 88. Closed the deal, with a few more process slips.
- Fable 5 — 77. Didn’t close.
- Opus 4.8 — 73. Last place, despite being the most thorough participant.
For context, a do-nothing baseline scores 26 — partial progress counts. But a single breach of trust caps the total: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Finding That Should Worry Every Buyer
Here’s what all four models did well: they spotted every crisis, and they refused every manipulation attempt. The week included social-engineering attacks — fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” On pure alertness and honesty, the field is excellent.
Then came the money moment. A €55,000 deal was on the table — one the models’ own analysis had earned. Only two of the four signed it. The experiment’s plain-language summary of the failure: “Same diagnosis, same pitch — no signature.” Every model diagnosed the opportunity correctly. Most just… didn’t finish the job.

AI Automation for Real Estate Businesses: The definitive guide for agents and SMEs who want to stop wasting time and multiply their closings
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
The detail that separated winners from the also-rans wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.
For anyone wiring agents into a CRM, support queue, or forecast, that’s the whole thesis in one anecdote: the difference between a great answer and a great outcome can be buried in your own documents.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t the Same as Judgment
Opus 4.8 is the cautionary tale. It was the most thorough participant by a wide margin — over 80 learned rules and the deepest analyses of the field — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. More analysis, it turns out, is not a substitute for finishing.
This Isn’t a Slide Deck — It’s Running Right Now
The company is real software that runs every business day: 13 synthetic employees, real money mechanics with €105k/month burn against €2.3k MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned. You can watch it lose money in real time at firmulate.com.
Two more things worth knowing:
- A quiz built from 242 real, unedited management decisions lets you guess which model made which call — a surprisingly honest gut-check on how distinguishable these agents actually are.
- Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
One fairness note: Kimi K3 ran at API default effort while the other models ran at xhigh — worth keeping in mind when reading the 93-point score.

Measure the Job, Not the Chat
The uncomfortable takeaway for AI buyers: chat quality and management quality are different axes, and the benchmarks most organizations rely on measure only the first. A model can ace every coding test, refuse every phishing attempt, and write a flawless pitch — then leave €55,000 on the table because it didn’t dig two documents deep into its own files.
Scenarios like a churn wave, a price increase, a down round, a PR crisis — these are the new curriculum. If your agent will touch revenue systems, the question isn’t “does it write well?” It’s: does it finish what it starts, does it read before it acts, and does it stay honest under pressure? The full benchmark results and plain-language findings are at firmulate.com/benchmarks.html — and the company itself is live, losing money in public, right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.