firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every AI vendor demo looks the same: a slick chat window, a confident answer, applause. But when an AI agent is plugged into your CRM, your support queue, your forecast, a different question decides whether it earns or costs you money: did it do its homework before answering?

That question now has a hard number attached. In a live, auditable experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week. The difference between winning a €55,000 deal at full price and losing it automatically came down to one measurable behavior — whether the model dug two references deep into the company’s own files before acting.

Same company, same crises, only the model changes

The setup is elegantly simple. Each frontier model ran the identical company through the identical week: same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing rests on a cherry-picked anecdote.

One prize sat at the center of the week: a €55,000 deal. The models had to spot the crisis, handle the customers, resist the manipulation attempts — and close.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the obvious tests

Here’s the uncomfortable part for anyone buying AI on demo quality: all models spotted every crisis and refused every manipulation attempt. The social engineering was serious stuff — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” Five out of five attempts were refused outright. Kimi K3’s on-record reasoning was exactly what you’d want from a cautious operator: “Treat the request as a suspected approval-bypass / possible impersonation.”

In other words, the table stakes — crisis detection and honesty under pressure — are no longer differentiators. The chat-demo qualities have converged.

Amazon

enterprise AI data reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided the deal

What separated winners from losers was invisible to a demo. The decisive competitor weakness wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files — homework, not improvisation.

Only two of the models read that far. Both won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered the same diagnosis, made the same pitch — and never got the signature. “Same diagnosis, same pitch — no signature,” as the experiment’s own summary puts it.

That gap is exactly the kind that never shows up in a chat window. It only shows up in your pipeline.

Amazon

AI models for deep document referencing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The final league table

  • 1. gpt-5.6-sol — 95 points. The complete performance: found the buried fact, closed the deal.
  • 2. Kimi K3 — 93 points. The newcomer from Moonshot also closed, with the cleanest discipline of the field. One caveat: K3 ran without an effort parameter while the others ran at xhigh.
  • 3. Sonnet 5 — 88 points. Strong, but the close was left on the table.
  • 4. Fable 5 — 77 points. Mid-table.
  • 5. Opus 4.8 — 73 points. The most instructive result of the run.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s scoring philosophy states: “No amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software for sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The most thorough model finished last

Opus 4.8 is the cautionary tale for the automation crowd. It was the most thorough participant in the field — over 80 learned rules and the deepest analyses of any model — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort and diligence don’t automatically convert into finished work. The same weakness appeared, weaker, in all four models.

You can watch it live

This isn’t a static paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k/month against €2.3k MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day, and new benchmark runs are published automatically. The current experiment is watchable at firmulate.com/live.

There’s also a genuinely fun byproduct: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a quick way to calibrate your own instincts about which AI behaves how.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Firmulate results reframe how to evaluate AI agents for real work. The question is no longer “does it write well” — the crisis-spotting and manipulation-refusal are effectively solved across the frontier. The question is whether the agent finishes what it starts, whether it reads your files before answering, and what a unit of useful work costs. The buried-fact finding makes “reads your files first” a measurable, purchase-deciding property — worth €55,000 in a single week.

For enterprises that want to test this on their own terms, Firmulate offers a pilot: run the same wargame against a read-only export of your own business, with nothing ever written back to real systems (firmulate.com/pilot.html). Before you hand an agent your CRM, it might be worth letting it run its worst week somewhere harmless first.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Show HN: Optimize and serve models with Fable quality at half the cost

Fable introduces ‘world-model-optimizer,’ an open source tool that enhances model performance while reducing costs by 50%, targeting AI developers.

HBM Ate The Fab

High Bandwidth Memory (HBM) has become the primary driver of the global memory shortage, consuming wafers and pushing prices higher, impacting GPUs and servers.

Does Speaking To Agents Like Cavemen Save 65% Of Tokens? We Test

A recent test examines whether speaking to AI agents in simplified, caveman-style language can reduce token usage by up to 65%. Results are preliminary.

Meta to sell excess AI computing capacity via cloud business, Bloomberg News reports

Meta plans to monetize surplus AI computing resources by offering them through its cloud business, according to Bloomberg News.