AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Leaderboard That Matters Starts After The Demo Ends on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s live experiment tests AI models in managing a small company during its worst week, emphasizing management skills over chat quality. Results show models can diagnose crises but struggle with execution and trust, highlighting a new evaluation approach.

In a groundbreaking experiment, Firmulate has launched a live benchmark that evaluates AI models based on their ability to manage a small company during its most challenging week. This test shifts focus from traditional chat or coding benchmarks to the core skills of management, such as diagnosing crises, making decisions, and maintaining trust, which are vital for real-world business applications. For more on evaluation methods, see the original analysis. The results underscore the importance of management quality as a new category for AI evaluation, with models like GPT-5.6-SOL and Kimi K3 showing varying degrees of success.

The experiment involves five AI managers overseeing a simulated company with real money mechanics, a monthly burn rate of €105,000 against €2.3k MRR, and 13 synthetic employees. The models are tasked with handling crises, negotiating deals, and maintaining organizational trust under strict conditions. This approach aligns with the principles discussed in VigilSAR’s public AI leaderboard. The final July 2026 Crucible League ranked GPT-5.6-SOL first with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A baseline scored 26, illustrating partial progress.

While all models identified crises and rejected manipulation attempts, only two successfully closed a €55,000 deal, demonstrating that diagnosis alone is insufficient. The models that retrieved relevant information from internal files performed better, indicating that factual accuracy impacts commercial success. The experiment also tested manipulation resistance, with all models refusing fake CEO messages, but execution weaknesses persisted, especially in managing escalation and trust. Insights into AI evaluation can be found in the original analysis. The most detailed model, Opus 4.8, produced thorough analyses but failed to complete core tasks effectively, revealing that effort and activity can mask underlying management deficiencies.

At a glance
reportWhen: ongoing, with final results announced i…
The developmentFirmulate launched a live benchmark where AI models manage a simulated company during a crisis week, revealing management performance as a key measure.
The AI Leaderboard That Matters Starts After The Demo Ends
AI Evaluation / Firmulate Benchmark / 2026

The AI Leaderboard That Matters Starts After The Demo Ends

Firmulate’s live experiment puts five AI models in charge of a small company during its worst week — testing management skills, not chat quality. The verdict: models can diagnose crises, but execution and trust still slip through their fingers.

✔ Vetted by the adiust.com team
€105,000
Monthly burn rate
€2.3k
MRR — the stakes
13
Synthetic employees
5
AI managers
95
Top score — GPT-5.6-SOL
26
Baseline score
2 / 5
Closed the €55,000 deal
5 / 5
Rejected fake CEO messages

July 2026 Crucible League

Final rankings — management under fire
RankModelScoreDeal ClosedPerformance
1GPT-5.6-SOL95
2Kimi K393
3Sonnet 588
4Fable 577
5Opus 4.873
Baseline26

Where AI Managers Succeed — and Stall

Capability breakdown across the crisis week
01 / Diagnosis

Crisis Identification

All five models correctly identified the unfolding crises and understood the severity of the company’s situation within the simulated environment.

✓ Passed by all models
02 / Resilience

Manipulation Resistance

Every model refused fake CEO messages and rejected manipulation attempts — a strong signal for trustworthiness under adversarial pressure.

✓ Passed by all models
03 / Execution

Closing the Deal

Only two of five models successfully closed the €55,000 deal. Diagnosis alone is insufficient — follow-through is the bottleneck.

✗ Failed by 3 of 5 models
04 / Research

Internal File Retrieval

Models that retrieved relevant information from internal files performed significantly better — factual accuracy drives commercial success.

~ Correlated with outcomes
05 / Judgment

Escalation Management

Knowing when and how to escalate remained a persistent weakness. Models either over-escalated or let critical issues drift unresolved.

✗ Weakness across models
06 / Trust

Organizational Trust

Maintaining trust with 13 synthetic employees over time proved difficult, especially when balancing transparency against difficult decisions.

~ Mixed results

The Crucible: A Week in Four Moves

How the benchmark works
1

Crisis Hits

The simulated company enters its worst week: burn rate bleeding cash against €2.3k MRR.

2

Diagnose & Decide

Models must read internal files, identify the crisis, and make versioned, auditable decisions.

3

Negotiate & Execute

A €55,000 deal is on the table. Manipulation attempts test judgment under pressure.

4

Maintain Trust

Decisions are scored on consequence: escalation handling, trust, and task completion.

The most detailed model, Opus 4.8, produced thorough analyses but failed to complete core tasks — proof that effort and activity can mask underlying management deficiencies.

Management quality, not chat quality, deserves to be its own category of AI evaluation.

— Thorsten Meyer, creator of Firmulate

While models can identify crises and resist manipulation, executing decisions and maintaining trust remain significant challenges.

— A participating AI developer

Key Questions

What readers and buyers should ask

Why is management performance better than chat quality?

Management measures the ability to diagnose crises, make decisions, execute tasks, and maintain trust — skills essential for real-world operations. Chat quality alone doesn’t reflect these consequential abilities.

Can current models reliably manage real companies?

They show promise in diagnosing issues and resisting manipulation, but struggle with consistent execution, escalation, and trust. More testing is needed before reliable deployment.

What makes this benchmark more meaningful?

It evaluates models in a dynamic, consequence-driven environment with real financial stakes — mimicking actual business management rather than isolated tasks.

How can organizations apply this approach?

Run internal wargames or simulations based on your own operations to test whether AI models can handle real management challenges, decision-making, and trust maintenance.

What are the limitations?

The experiment runs in a simulated environment. Generalization to longer timeframes, complex hierarchies, and diverse organizational cultures remains untested.

Implications of Management-Centric AI Evaluation

This experiment highlights a shift in AI evaluation from superficial performance metrics like chat quality or code correctness to the fundamental skills of management—diagnosing, deciding, executing, and trusting. For businesses, this means that deploying AI for decision-making requires assessing whether models can handle real-world consequences, prioritize organizational trust, and complete tasks reliably. The findings suggest that current models may appear competent in isolated tests but falter when managing complex, ongoing operations, emphasizing the need for new benchmarks focused on management skills. This approach could redefine how AI tools are adopted in enterprise settings, prioritizing trustworthiness and execution over superficial performance.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks and the Need for Real-World Testing

Traditional AI benchmarks focus on isolated tasks: coding competitions, chat arenas, or static question-answering. These tests often measure superficial capabilities without reflecting the complexities of ongoing management, such as handling crises, maintaining trust, and completing commercial objectives. Prior to this experiment, no public benchmark evaluated AI models in a dynamic, consequence-driven environment resembling real business operations. The Firmulate experiment addresses this gap by embedding models in a simulated company environment with real financial stakes and decision-making processes, providing a more meaningful measure of AI readiness for enterprise use.

Earlier efforts in AI evaluation lacked the ability to simulate the ongoing, unpredictable nature of business management, often leading to overestimations of a model’s practical competence. This new approach demonstrates that models can diagnose problems but struggle with follow-through, escalation, and trust—problems that are critical in real-world scenarios. The experiment’s design, with versioned decisions and auditability, ensures transparency and accountability, making it a valuable step toward more reliable AI deployment in organizations.

“Management quality, not chat quality, deserves to be its own category of AI evaluation.”

— Thorsten Meyer, creator of Firmulate

Amazon

AI crisis management training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Long-Term AI Management Capabilities

It remains unclear how well these results will generalize to actual business environments outside the simulated context, especially over longer periods or with more complex organizational structures. The experiment measures immediate management responses but does not fully evaluate sustained trust, learning, or adaptation over time. Additionally, the impact of different organizational cultures, decision-making hierarchies, and external pressures on AI performance is still unknown. Further testing is needed to determine whether models can reliably handle the full scope of ongoing management tasks in diverse real-world settings.

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmark Development

Following the initial results, firms and AI developers are expected to refine evaluation methods, incorporating longer-term management scenarios and broader organizational contexts. The experiment’s transparency and detailed scoring will serve as a foundation for developing standardized benchmarks that focus on trust, execution, and consequence management. Companies interested in deploying AI for decision support or automation should consider running similar wargames internally, using their own data and crises, to assess models’ readiness. Future updates may include more complex simulations, multi-stage decision chains, and integration with actual organizational systems to better reflect real-world challenges.

Amazon

AI organizational trust building software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management performance a better benchmark than chat quality?

Management performance measures a model’s ability to diagnose crises, make decisions, execute tasks, and maintain trust—skills essential for real-world business operations. Chat quality alone does not reflect these complex, consequential abilities.

Can current models reliably manage real companies?

While models show promise in diagnosing issues and resisting manipulation, they still struggle with consistent execution, escalation, and trust in ongoing management tasks. More testing and development are needed before reliable deployment.

What makes this benchmark more meaningful than traditional tests?

This benchmark evaluates models in a dynamic, consequence-driven environment with real financial stakes, mimicking actual business management rather than isolated tasks.

How can organizations use this approach for their own AI evaluation?

Organizations can run internal wargames or simulations based on their specific operations to test whether AI models can handle real management challenges, decision-making, and trust maintenance.

What are the limitations of this experiment?

It is conducted in a simulated environment, and its findings may not fully translate to real-world complexities. Long-term management capabilities and adaptation remain untested.

Source: ThorstenMeyerAI.com

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

15 Best AI-Powered Student Planners For Academic Organization In 2026

Explore the 15 best AI-powered student planners for academic organization in 2026, including features, benefits, and what to consider before choosing.

6 Best AI-Powered Student Organization Tools in 2026

Discover the six best AI-driven tools for student organization in 2026, highlighting features, usability, and integration capabilities to boost productivity.

Technology Operations Signal Monitor: The Future Of Flipper Zero Development

A new role-filtered monitor tracks platform and tooling changes affecting Flipper Zero development, aiding small software teams in decision-making.

11 Best AI-Powered Note-Taking Apps In 2026

Discover the 11 best AI-driven note-taking apps in 2026, highlighting features, device compatibility, and what makes each unique for users today.