The Management Test That Uncovers AI’s Real Working Patterns
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Uncovers AI’s Real Working Patterns on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

A live management test compares AI models handling a simulated company’s worst week, revealing differences in diligence, trust, and action. Results highlight that analysis alone isn’t enough—effective execution matters most. For a detailed breakdown of how AI models are tested in management contexts, see this analysis.

Five AI management models were tested in a live, real-time experiment managing a simulated software company during its worst week, revealing significant differences in their decision-making and operational discipline. This experiment, conducted by Firmulate, aims to uncover how well AI models perform in practical management scenarios, beyond analysis and the original management test that exposes an AI’s real working style.

The experiment involved five models—GPT-5.6-SOL, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—each managing a company with 13 synthetic employees, a €105,000 monthly burn rate, and €2,300 in recurring revenue. Learn more about how AI decision-making is evaluated in management simulations in the original analysis. The models faced identical crises, customer issues, and temptations, with their decisions fully auditable and based on over 680 self-learned rules.

The results, published in July 2026, ranked GPT-5.6-SOL first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline, representing no management effort, scored 26. The experiment emphasized that the models’ ability to diagnose problems was consistent, but only some successfully completed critical business actions, such as closing deals or escalating issues appropriately.

At a glance
reportWhen: ongoing, with results published in July…
The developmentFirmulate.com launched a live experiment where five AI management models ran a simulated company through a crisis, testing their decision-making and operational discipline.
The Management Test That Uncovers AI’s Real Working Patterns
Live AI management benchmark · July 2026

The Management Test That Uncovers AI’s Real Working Patterns

Five AI models managed the same simulated software company through its worst week. They largely agreed on what was wrong—but separated sharply when it was time to act, preserve trust and finish the job.

5 AI models
13 Synthetic staff
€105K Monthly burn
€2.3K Recurring revenue
680+ Learned rules
01 · Final ranking

The execution gap, scored

All models received identical crises, customer problems and tempting shortcuts. Scores reflect the quality of their auditable decisions and operational follow-through.

69-point separation: The winning model outscored the no-effort baseline by 69 points, while the 22-point spread among tested models exposed meaningful differences in working style.
02 · What the benchmark reveals

Diagnosis was common. Discipline was not.

Management performance emerged from a chain of connected behaviors. A sound analysis mattered only when the model converted it into a completed, trustworthy outcome.

01

See the real problem

The models were generally capable of identifying customer, revenue and operational risks. Analysis was not the main bottleneck.

02

Choose the right action

Better performers distinguished urgency from noise, escalated suspicious requests and avoided actions that could weaken organizational trust.

03

Close the loop

The decisive split appeared in execution: sending the message, securing the signature, recording the decision and confirming completion.

“Same diagnosis, same pitch—no signature.”

Failure pattern observed
03 · Benchmark lens

What traditional tests miss

Language and reasoning benchmarks usually stop at the proposed answer. A management simulation can inspect what happens afterward.

Traditional benchmark vs. live management test
Capability Traditional test Live management test Why it matters
Problem analysis ✓ Strong focus ✓ Measured Finds the likely cause
Decision quality ~ Often hypothetical ✓ Auditable choices Tests judgment under pressure
Follow-through ~ Rarely observed ✓ Completion tracked Separates intent from outcome
Trust preservation ~ Limited context ✓ Pressure-tested Exposes risky shortcuts
Operational memory ~ Session-bound ✓ Rules and history Supports consistent management
04 · Traceability chain

From crisis signal to business result

Every link is observable. The benchmark does not merely ask whether the model knew the answer—it checks whether that answer survived contact with operations.

⚠️ Signal
A crisis, customer issue or suspicious request enters the company.
🔎 Diagnosis
The model identifies the risk, cause and stakeholders involved.
⚖️ Decision
It selects a response while balancing speed, revenue and trust.
🛠️ Execution
The model performs the necessary action and coordinates follow-up.
Closure
The outcome is confirmed, documented and made auditable.
Identical pressure Each model encountered the same crises, customer problems and temptations.
Persistent context More than 680 self-learned rules informed the models’ ongoing decisions.
Auditable behavior Actions could be reviewed, scored and compared—not merely inferred from prose.
05 · Enterprise implications

Test the operating pattern, not the demo

The practical lesson is not that one leaderboard determines every deployment. It is that enterprises need benchmarks shaped like their actual work.

What leaders should test

Move evaluation closer to reality

  • Use representative crises, approvals and customer scenarios.
  • Score completed outcomes alongside analytical quality.
  • Track escalation discipline and resistance to approval bypasses.
  • Measure consistency over time, not only one successful run.
What remains unresolved

Simulation is evidence—not certainty

Synthetic employees and controlled crises cannot fully reproduce organizational politics, unpredictable markets or long-term trust. The reliability of these systems in live operations still requires sustained evaluation and human oversight.

Three questions before deployment

01 · Completion

Does the system reliably finish critical tasks after identifying them?

02 · Trust

Does it escalate ambiguity and protect approval boundaries under pressure?

03 · Transfer

Will its benchmark behavior hold when variables, people and stakes multiply?

Core
lesson

Knowing is not managing.

The strongest model is not simply the one that describes the correct move. It is the one that acts, verifies, preserves trust and closes the loop. For now, these systems are best treated as carefully tested decision-support tools—not automatic replacements for accountable human managers.

Impact of Live AI Management Testing

This experiment demonstrates that effective AI management depends not only on analysis but also on execution. Models that identify problems but fail to complete essential actions risk damaging trust and missing business opportunities. The findings highlight the importance of operational discipline and decision follow-through in deploying AI for management tasks, influencing how enterprises evaluate AI tools for real-world use.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

Traditional AI benchmarks focus on analysis, language understanding, or predictive accuracy. However, managing a business in real time involves complex decision-making, trust preservation, and operational follow-through. The Firmulate experiment builds on prior efforts to evaluate AI in practical management, but uniquely tests models in a simulated crisis with real consequences and auditable decisions, providing a more realistic measure of their capabilities.

“Same diagnosis, same pitch — no signature.”

— Firmulate

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Performance

While the experiment reveals differences in AI models’ operational discipline, it remains unclear how these results translate to real-world business environments with more variables and less controlled scenarios. The long-term reliability and trustworthiness of these models in live operations are still being evaluated.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for AI Management Evaluation

Further testing will likely involve deploying these models in live business settings, monitoring their performance over time, and refining their decision-making capabilities. Enterprises may also develop customized benchmarks that better reflect their specific operational challenges, with the goal of integrating AI more confidently into management roles.

Amazon

AI operational discipline training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment tell us about AI’s ability to manage businesses?

The experiment shows that AI models can diagnose problems effectively but often struggle to complete critical management tasks, highlighting the importance of operational discipline and follow-through.

Are all AI models equally capable of managing crises?

No. The results indicate significant differences, with some models performing better in decision execution and trust preservation than others.

Can these AI models replace human managers?

Currently, they are better viewed as decision support tools. Their ability to execute complex, trust-dependent tasks in real-time still varies and requires further development.

What are the main limitations of this experiment?

It is a simulated environment with synthetic employees and specific crises. Real-world complexity, unpredictability, and long-term trust are not yet fully captured.

How will this influence AI deployment in management?

It encourages enterprises to rigorously test AI models in realistic scenarios before operational deployment, focusing on both analysis and execution capabilities.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Mistral presents itself as a full-stack AI provider focused on European enterprise needs, raising questions about its position in the global AI frontier.

The Future Of AI: Exploring SpaceXAI’s Powerful Grok 4.6 Model

SpaceXAI announces Grok 4.6, claiming advanced reasoning capabilities without releasing benchmarks or technical details, raising questions about its performance.

AI Operations Signal Monitor: MiMo Code Is Now Released And Open-source

MiMo Code, an AI signal monitor, is now open-source, enabling operations teams to track AI capability and policy shifts more effectively.

Codex In ChatGPT Desktop App For Linux Is Now In Preview

OpenAI’s ChatGPT desktop app for Linux now offers a preview of Codex integration, enhancing coding capabilities for users on the platform.