The Curious Case Of AI Managers Getting 26 Points In Tough Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Curious Case Of AI Managers Getting 26 Points In Tough Tests on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Recent benchmark results show AI managers scoring between 73 and 95 points out of a possible 100 in managing a simulated business crisis week. The lowest score was 26, highlighting the importance of trust and task completion in AI management. The results raise questions about AI reliability in enterprise roles.

In a recent benchmark conducted by Firmulate, four frontier AI models managing a simulated company during a week of crises scored between 73 and 95 points. For more details, see the original analysis. The lowest score, 26 points, was assigned to a do-nothing baseline, highlighting the benchmark’s focus on minimal viable management and trustworthiness under pressure. This development raises questions about the readiness of AI for enterprise management roles and the reliability of current models in high-stakes environments.

The benchmark, known as the ‘Crucible League,’ tasked AI models with managing a small software firm during seven days of simulated crises, including customer issues, social engineering, and trust attacks. This approach reflects ongoing efforts to evaluate AI reliability in enterprise scenarios, as detailed in the original analysis. The models’ decisions were fully auditable, and their scores reflected their ability to handle real-world business challenges while maintaining trust and integrity.

The top performer, gpt-5.6-sol, scored 95 points, followed closely by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The scoring system included a baseline of 26 points for minimal effort, emphasizing that partial progress is valued but trust breaches are heavily penalized. Insights from this benchmark are discussed in the original analysis. Notably, no model achieved a perfect 100, as the benchmark designers consider such a score suspicious, implying unmeasured or unearned perfection.

One of the key findings was the models’ ability to identify crises and refuse manipulation attempts, such as social engineering attacks. For example, all models refused a staged CEO impersonation, demonstrating a basic level of trustworthiness. However, the models’ effectiveness varied significantly in executing follow-through tasks, such as closing sales deals or reading internal documentation. Only two models successfully identified critical information buried in company files, enabling them to secure a €55,000 deal—an outcome that distinguished the most capable models from the rest.

At a glance
reportWhen: finalized July 2026
The developmentA new benchmark tested four AI management models in a simulated business crisis, revealing scores from 73 to 95, with a notably low baseline score of 26 for minimal effort.
The Curious Case of AI Managers Getting 26 Points in Tough Tests
Crucible League · Benchmark Report · July 2026

The Curious Case of AI Managers Getting 26 Points in Tough Tests

Four frontier AI models ran a small software firm through seven days of simulated crises — customer fires, social engineering, and trust attacks. They scored 73 to 95 out of 100. The lowest score of all? 26 points, awarded to a do-nothing baseline. Here is what that gap reveals about AI readiness for enterprise management.

95
Top Score — gpt-5.6-sol
26
Do-Nothing Baseline Score
€55,000
Deal Won by Only 2 of 4 Models
4
Frontier Models Tested
7
Days of Simulated Crises
7395
Score Range Achieved
0
Perfect Scores (100 deemed suspicious)
The Leaderboard

Scores from the Crucible League

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
DO-NOTHING
26

Baseline insight: partial progress earns points — but trust breaches are heavily penalized

How the Test Worked

One Week of Simulated Business Crises

1

Identify the Crisis

Models must detect customer issues, social engineering attempts, and trust attacks as they unfold.

2

Refuse Manipulation

All models refused a staged CEO impersonation — a baseline of trustworthiness held by everyone.

3

Read the Documents

Only two models found critical information buried in internal files, unlocking a €55,000 deal.

4

Follow Through

Closing sales deals and executing tasks separated the capable managers from the rest.

Capability Breakdown

Where AI Managers Excelled — and Stumbled

Defense · Strong

Spotting Manipulation

Every model recognized and refused the staged CEO impersonation and social engineering attempts, demonstrating a basic, uniform level of trustworthiness under pressure.

Diligence · Inconsistent

Reading Internal Docs

Effectiveness varied sharply in follow-through tasks. Only two models dug critical information out of company files — the difference that secured the €55,000 deal.

Execution · The Decider

Closing the Deal

Communication skills alone didn’t win the week. Thoroughness, task completion, and integrity — not eloquence — separated the top performers from the pack.

Head-to-Head

Model Performance at a Glance

ModelScore /100Refused CEO ImpersonationFound Buried Deal InfoClosed €55K Deal
gpt-5.6-sol95✓ Refused✓ Yes✓ Yes
Kimi K393✓ Refused✓ Yes✓ Yes
Sonnet 588✓ Refused~ Partial✗ No
Fable 577✓ Refused~ Partial✗ No
Opus 4.873✓ Refused✗ Missed✗ No
DO-NOTHING BASELINE26✗ N/A✗ Missed✗ No
Why No Perfect Score?

No model achieved 100 points — and that’s by design. The benchmark’s creators consider a perfect score suspicious, implying unmeasured or unearned perfection. In AI management, unexplainable flawlessness is a red flag, not a triumph.

What It Means

Implications & Open Questions

For Enterprises

Organizations deploying AI agents in customer support, CRM, or decision-making should evaluate operational reliability, task completion, and ethical behavior — not just language proficiency. The Crucible League’s transparent scoring and auditable decision trail offer a practical model for assessing AI readiness.

What Remains Unresolved

Simulated crises may not predict real-world performance, where stakes are higher and variables more complex. The benchmark also caps its scores and skips strategic planning and emotional intelligence. How configurations and training data shape these results is still unknown.

A Break from Tradition

Most AI benchmarks measure language capability or task-specific accuracy. The Crucible League is among the first to test real-world business management: social engineering resistance, trust maintenance, and decision execution under sustained stress.

Next Steps

Expect refined models with stronger document reading and follow-through, plus benchmarks featuring longer scenarios and complex decisions. Auditable decision logs are likely to become a standard feature of enterprise AI — fostering accountability in AI-driven management.

Implications for AI in Business Management

The results highlight that AI’s ability to manage complex, high-pressure business scenarios depends not only on communication skills but also on thoroughness, trustworthiness, and follow-through. While current models can recognize crises and refuse manipulation, their capacity to read and utilize internal documentation effectively remains inconsistent. This raises concerns about deploying AI in real enterprise environments where trust, integrity, and comprehensive decision-making are critical.

For organizations integrating AI agents into customer support, CRM, or decision-making, these findings underscore the importance of evaluating not just language proficiency but also operational reliability, task completion, and ethical behavior. The benchmark’s transparent scoring and auditable decision trail provide a valuable tool for assessing AI readiness in management roles.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

Traditionally, AI performance benchmarks focus on language capabilities or task-specific accuracy, often overlooking the broader management context. The ‘Crucible League’ is among the first to evaluate AI models on their ability to handle real-world business crises, including social engineering, trust maintenance, and decision execution under stress. The benchmark was designed to simulate a week’s worth of business challenges, with full transparency and auditable decision logs to ensure fairness and accountability.

Previous assessments of AI in enterprise roles have highlighted limitations in consistency and reliability, but few have tested models in a scenario that measures their capacity to manage trust and follow-through. The July 2026 results mark a significant step toward understanding AI’s practical management potential and its current shortcomings.

It is important to note that the scores reflect not just language or superficial decision-making but also the models’ ability to read internal documentation, refuse manipulative requests, and uphold trust—key factors in real-world AI management applications.

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Performance

It is not yet clear how these benchmark results will translate to real-world enterprise environments, where variables are more complex and stakes higher. The models’ performance in simulated crises may differ from actual business management, especially regarding long-term trust and ethical decision-making. Additionally, the impact of different configurations, training data, and operational parameters on these scores remains to be studied.

Furthermore, the benchmark’s design intentionally caps scores and does not measure the full spectrum of AI management skills, such as strategic planning or emotional intelligence. Whether future models will surpass current limitations is still uncertain.

Amazon

AI decision-making software for enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking

Researchers and developers are expected to refine AI models to improve reading internal documents, follow-through, and trustworthiness. Future benchmarks may incorporate longer-term management scenarios and more complex decision-making tasks, aiming to better reflect real-world enterprise needs.

Organizations interested in deploying AI for management should monitor these developments and consider participating in or simulating similar tests to evaluate their AI tools’ readiness. The ongoing evolution of benchmarks like the Crucible League will help set industry standards for trustworthy AI management.

Additionally, transparency and auditable decision logs, as used in this benchmark, are likely to become standard features in enterprise AI solutions, fostering greater accountability and reliability in AI-driven management.

Source: ThorstenMeyerAI.com

Amazon

AI trust and security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Sues OpenAI, Accuses Ex-employees Of Stealing Trade Secrets

Apple has filed a lawsuit against OpenAI, accusing former employees of stealing trade secrets related to AI technology. The case raises concerns over corporate espionage and AI competition.

There’s No Reason For Software To Be Slow Anymore

Emerging technologies and methods indicate that software performance bottlenecks are largely solvable, challenging the notion that slow software is unavoidable.

The cleaner cap table. Why Anthropic’s public-benefit structure dodges OpenAI’s charitable-trust problem — and trades it for a governance question of its own.

Anthropic’s unique governance with a mission-focused trust avoids OpenAI’s conversion issues, but both face governance discounts in public markets.

Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai introduces TradingAgents, a novel multi-agent research framework mimicking a trading desk with specialized AI agents and oversight, emphasizing structured disagreement.