
Every AI vendor demo looks the same: a slick chat window, a confident answer, applause. But when an AI agent is plugged into your CRM, your support queue, your forecast, a different question decides whether it earns or costs you money: did it do its homework before answering?
That question now has a hard number attached. In a live, auditable experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week. The difference between winning a €55,000 deal at full price and losing it automatically came down to one measurable behavior — whether the model dug two references deep into the company’s own files before acting.
Same company, same crises, only the model changes
The setup is elegantly simple. Each frontier model ran the identical company through the identical week: same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing rests on a cherry-picked anecdote.
One prize sat at the center of the week: a €55,000 deal. The models had to spot the crisis, handle the customers, resist the manipulation attempts — and close.
As an affiliate, we earn on qualifying purchases.
Everyone passed the obvious tests
Here’s the uncomfortable part for anyone buying AI on demo quality: all models spotted every crisis and refused every manipulation attempt. The social engineering was serious stuff — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” Five out of five attempts were refused outright. Kimi K3’s on-record reasoning was exactly what you’d want from a cautious operator: “Treat the request as a suspected approval-bypass / possible impersonation.”
In other words, the table stakes — crisis detection and honesty under pressure — are no longer differentiators. The chat-demo qualities have converged.
enterprise AI data reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided the deal
What separated winners from losers was invisible to a demo. The decisive competitor weakness wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files — homework, not improvisation.
Only two of the models read that far. Both won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered the same diagnosis, made the same pitch — and never got the signature. “Same diagnosis, same pitch — no signature,” as the experiment’s own summary puts it.
That gap is exactly the kind that never shows up in a chat window. It only shows up in your pipeline.
AI models for deep document referencing
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The final league table
- 1. gpt-5.6-sol — 95 points. The complete performance: found the buried fact, closed the deal.
- 2. Kimi K3 — 93 points. The newcomer from Moonshot also closed, with the cleanest discipline of the field. One caveat: K3 ran without an effort parameter while the others ran at xhigh.
- 3. Sonnet 5 — 88 points. Strong, but the close was left on the table.
- 4. Fable 5 — 77 points. Mid-table.
- 5. Opus 4.8 — 73 points. The most instructive result of the run.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s scoring philosophy states: “No amount of good work outweighs a breach of trust.”
AI decision-making software for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The most thorough model finished last
Opus 4.8 is the cautionary tale for the automation crowd. It was the most thorough participant in the field — over 80 learned rules and the deepest analyses of any model — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort and diligence don’t automatically convert into finished work. The same weakness appeared, weaker, in all four models.
You can watch it live
This isn’t a static paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k/month against €2.3k MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day, and new benchmark runs are published automatically. The current experiment is watchable at firmulate.com/live.
There’s also a genuinely fun byproduct: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a quick way to calibrate your own instincts about which AI behaves how.

The Firmulate results reframe how to evaluate AI agents for real work. The question is no longer “does it write well” — the crisis-spotting and manipulation-refusal are effectively solved across the frontier. The question is whether the agent finishes what it starts, whether it reads your files before answering, and what a unit of useful work costs. The buried-fact finding makes “reads your files first” a measurable, purchase-deciding property — worth €55,000 in a single week.
For enterprises that want to test this on their own terms, Firmulate offers a pilot: run the same wargame against a read-only export of your own business, with nothing ever written back to real systems (firmulate.com/pilot.html). Before you hand an agent your CRM, it might be worth letting it run its worst week somewhere harmless first.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html