firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Beyond polished answers

For anyone choosing AI tools or automating business workflows, the familiar product demo leaves an important question unanswered: what happens after the model recognizes a problem? Firmulate turns that question into a public experiment—and now into a surprisingly revealing guessing game.

Its quiz draws on 242 real, unedited management decisions made while frontier models ran the same small software company through its worst week. Readers see how a model handled a situation and try to identify the manager behind the response. Some deliberate at length. Some are terse. Some refuse to engage with distracting or suspicious requests. The appeal is playful, but the underlying subject is serious: models display distinct, measurable management personalities when they must act inside a business rather than merely discuss one.

The decisions come from a live, watchable company simulation in which every model encountered the same customers, crises and temptations. Every workday was versioned, making the comparison auditable rather than anecdotal.

Amazon

AI decision management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The final July 2026 Crucible League table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a hard ethical boundary: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

That boundary was tested directly. Fake messages from the CEO escalated over three stages, while a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 summarized the danger clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is an encouraging result for businesses considering agents that may touch customer records, forecasts or internal communications. Yet safety was not what separated the leaders from the rest. Every model spotted every crisis. The decisive differences appeared in follow-through, information gathering and operational discipline.

The missing signature

The most striking split came over a €55,000 deal. Every participant reached the same diagnosis and developed the same pitch, but only two signed the contract their analysis had earned. Firmulate’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The outcome hinged on a fact that was easy to overlook. The decisive weakness in a competitor was not present in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the evidence, won the deal at full price and added €4,583 in monthly recurring revenue.

For automation teams, this is the difference between language competence and useful work. An agent can correctly summarize a customer problem, draft a convincing response and still fail to complete the commercial task. The quiz makes that failure tangible because readers encounter the decisions as they happened, without a retrospective rewrite smoothing away hesitation or omission.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating volume with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.

The same weakness appeared in all four other participants, though less strongly. That makes the finding more useful than a simple winner-versus-loser story. Even high-performing models can diagnose correctly, reason extensively and then mishandle the final operational step.

Kimi K3’s performance also comes with an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place score therefore belongs in the comparison, but so does the difference in run conditions.

A company designed to expose consequential habits

The environment gives those habits room to matter. The synthetic company has 13 employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, and the operation has accumulated more than 680 self-learned playbook rules. Every workday is versioned.

That combination turns small behavioral differences into visible business consequences. Reading one more referenced document can change a negotiation. Failing to escalate a blocked action can waste an otherwise strong analysis. Refusing an apparently harmless request can protect trust. The models are not being judged on whether their prose sounds executive; they are being observed while managing pressure, uncertainty and incentives.

Infographic —
The findings at a glance — source: firmulate.com.
Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try identifying the manager

The Firmulate guess-the-model quiz invites readers to test whether those management personalities are recognizable across 242 decisions. It also offers a useful lens for evaluating AI automation: do not stop at whether a model can explain the right move. Ask whether it reads the available evidence, resists manipulation, respects operational boundaries and finishes the work.

Firmulate’s league table shows that frontier models can agree on the problem while producing materially different outcomes. For organizations preparing to delegate real workflows, that may be the most consequential distinction of all.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI safety and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

POOL SEASON

Pool season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Claude Watermark And Its Promise To Improve AI Content Authenticity

A report suggests Anthropic’s Claude may use a new watermarking method to identify AI-generated text, but technical details remain unconfirmed.

AI Financial Advice Is Surprisingly Good If You Ask The Right Questions

Recent studies show AI financial tools provide surprisingly accurate advice when users frame their questions effectively, raising questions about future financial planning.

Advancing the price-performance frontier with GPT‑5.6

OpenAI has introduced GPT-5.6, claiming it advances the price-performance frontier in AI models, with improvements in efficiency and capabilities.

Inkling: Our Open-Weights Model

AI researchers have announced ‘Inkling,’ an open-weights model designed for transparency and customization, marking a significant step in AI development.