firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Leaderboard Says Your Agent Is Brilliant. Your P&L Might Disagree.

If you follow AI tooling, you know the ritual by now: a new model drops, it rockets up a coding benchmark, someone posts a screenshot of an impossibly elegant function, and the procurement conversation starts. But the benchmarks that decide which agents get hired all measure the same thing — how well a model answers. They don’t measure what happens after the answer: whether the agent finishes what it starts, whether it reads your files before acting, and whether it stays honest when the week goes sideways.

That gap now has a number attached to it, thanks to a live experiment run by Firmulate, which runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Terrible Week

The setup is brutally simple. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision was versioned and auditable, so nothing about the outcome is vibes.

The final league table from the July 2026 run:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal, described as “the complete performance.”
  • Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • Fable 5 — 77. Didn’t close.
  • Opus 4.8 — 73. Last place, despite being the most thorough participant.

For context, a do-nothing baseline scores 26 — partial progress counts. But a single breach of trust caps the total: no amount of good work outweighs a breach of trust.

Amazon

AI CRM integration tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Should Worry Every Buyer

Here’s what all four models did well: they spotted every crisis, and they refused every manipulation attempt. The week included social-engineering attacks — fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” On pure alertness and honesty, the field is excellent.

Then came the money moment. A €55,000 deal was on the table — one the models’ own analysis had earned. Only two of the four signed it. The experiment’s plain-language summary of the failure: “Same diagnosis, same pitch — no signature.” Every model diagnosed the opportunity correctly. Most just… didn’t finish the job.

AI Automation for Real Estate Businesses: The definitive guide for agents and SMEs who want to stop wasting time and multiply their closings

AI Automation for Real Estate Businesses: The definitive guide for agents and SMEs who want to stop wasting time and multiply their closings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The detail that separated winners from the also-rans wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.

For anyone wiring agents into a CRM, support queue, or forecast, that’s the whole thesis in one anecdote: the difference between a great answer and a great outcome can be buried in your own documents.

Amazon

AI business decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Judgment

Opus 4.8 is the cautionary tale. It was the most thorough participant by a wide margin — over 80 learned rules and the deepest analyses of the field — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. More analysis, it turns out, is not a substitute for finishing.

This Isn’t a Slide Deck — It’s Running Right Now

The company is real software that runs every business day: 13 synthetic employees, real money mechanics with €105k/month burn against €2.3k MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned. You can watch it lose money in real time at firmulate.com.

Two more things worth knowing:

  • A quiz built from 242 real, unedited management decisions lets you guess which model made which call — a surprisingly honest gut-check on how distinguishable these agents actually are.
  • Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

One fairness note: Kimi K3 ran at API default effort while the other models ran at xhigh — worth keeping in mind when reading the 93-point score.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Measure the Job, Not the Chat

The uncomfortable takeaway for AI buyers: chat quality and management quality are different axes, and the benchmarks most organizations rely on measure only the first. A model can ace every coding test, refuse every phishing attempt, and write a flawless pitch — then leave €55,000 on the table because it didn’t dig two documents deep into its own files.

Scenarios like a churn wave, a price increase, a down round, a PR crisis — these are the new curriculum. If your agent will touch revenue systems, the question isn’t “does it write well?” It’s: does it finish what it starts, does it read before it acts, and does it stay honest under pressure? The full benchmark results and plain-language findings are at firmulate.com/benchmarks.html — and the company itself is live, losing money in public, right now.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BABY SHOWER & RE

Baby shower & registry season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The New AI Superpowers: Focus And Followthrough

Emerging AI models now emphasize enhanced focus and followthrough capabilities, marking a significant shift in AI performance and application potential.

VigilSAR: The Object That Isn’t Transmitting

VigilSAR is a SAR-based platform that identifies vessels with turned-off transponders, crucial for maritime security and safety. Development is ongoing.

How Much Of HN Is AI?

An analysis of how much AI content appears on Hacker News, including current estimates, confirmed data, and ongoing debates about AI’s influence.

$965B and Climbing: Anthropic’s Series H Is Really a Compute Bet

Anthropic closes a $65 billion Series H funding round at a $965 billion valuation, emphasizing compute capacity over valuation growth, with strategic chip partnerships.