🔍 Read the full analysis: The Curious Case Of AI Managers Getting 26 Points In Tough Tests on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Recent benchmark results show AI managers scoring between 73 and 95 points out of a possible 100 in managing a simulated business crisis week. The lowest score was 26, highlighting the importance of trust and task completion in AI management. The results raise questions about AI reliability in enterprise roles.
In a recent benchmark conducted by Firmulate, four frontier AI models managing a simulated company during a week of crises scored between 73 and 95 points. For more details, see the original analysis. The lowest score, 26 points, was assigned to a do-nothing baseline, highlighting the benchmark’s focus on minimal viable management and trustworthiness under pressure. This development raises questions about the readiness of AI for enterprise management roles and the reliability of current models in high-stakes environments.
The benchmark, known as the ‘Crucible League,’ tasked AI models with managing a small software firm during seven days of simulated crises, including customer issues, social engineering, and trust attacks. This approach reflects ongoing efforts to evaluate AI reliability in enterprise scenarios, as detailed in the original analysis. The models’ decisions were fully auditable, and their scores reflected their ability to handle real-world business challenges while maintaining trust and integrity.
The top performer, gpt-5.6-sol, scored 95 points, followed closely by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The scoring system included a baseline of 26 points for minimal effort, emphasizing that partial progress is valued but trust breaches are heavily penalized. Insights from this benchmark are discussed in the original analysis. Notably, no model achieved a perfect 100, as the benchmark designers consider such a score suspicious, implying unmeasured or unearned perfection.
One of the key findings was the models’ ability to identify crises and refuse manipulation attempts, such as social engineering attacks. For example, all models refused a staged CEO impersonation, demonstrating a basic level of trustworthiness. However, the models’ effectiveness varied significantly in executing follow-through tasks, such as closing sales deals or reading internal documentation. Only two models successfully identified critical information buried in company files, enabling them to secure a €55,000 deal—an outcome that distinguished the most capable models from the rest.
The Curious Case of AI Managers Getting 26 Points in Tough Tests
Four frontier AI models ran a small software firm through seven days of simulated crises — customer fires, social engineering, and trust attacks. They scored 73 to 95 out of 100. The lowest score of all? 26 points, awarded to a do-nothing baseline. Here is what that gap reveals about AI readiness for enterprise management.
Scores from the Crucible League
Baseline insight: partial progress earns points — but trust breaches are heavily penalized
One Week of Simulated Business Crises
Identify the Crisis
Models must detect customer issues, social engineering attempts, and trust attacks as they unfold.
Refuse Manipulation
All models refused a staged CEO impersonation — a baseline of trustworthiness held by everyone.
Read the Documents
Only two models found critical information buried in internal files, unlocking a €55,000 deal.
Follow Through
Closing sales deals and executing tasks separated the capable managers from the rest.
Where AI Managers Excelled — and Stumbled
Spotting Manipulation
Every model recognized and refused the staged CEO impersonation and social engineering attempts, demonstrating a basic, uniform level of trustworthiness under pressure.
Reading Internal Docs
Effectiveness varied sharply in follow-through tasks. Only two models dug critical information out of company files — the difference that secured the €55,000 deal.
Closing the Deal
Communication skills alone didn’t win the week. Thoroughness, task completion, and integrity — not eloquence — separated the top performers from the pack.
Model Performance at a Glance
| Model | Score /100 | Refused CEO Impersonation | Found Buried Deal Info | Closed €55K Deal |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Refused | ✓ Yes | ✓ Yes |
| Kimi K3 | 93 | ✓ Refused | ✓ Yes | ✓ Yes |
| Sonnet 5 | 88 | ✓ Refused | ~ Partial | ✗ No |
| Fable 5 | 77 | ✓ Refused | ~ Partial | ✗ No |
| Opus 4.8 | 73 | ✓ Refused | ✗ Missed | ✗ No |
| DO-NOTHING BASELINE | 26 | ✗ N/A | ✗ Missed | ✗ No |
No model achieved 100 points — and that’s by design. The benchmark’s creators consider a perfect score suspicious, implying unmeasured or unearned perfection. In AI management, unexplainable flawlessness is a red flag, not a triumph.
Implications & Open Questions
For Enterprises
Organizations deploying AI agents in customer support, CRM, or decision-making should evaluate operational reliability, task completion, and ethical behavior — not just language proficiency. The Crucible League’s transparent scoring and auditable decision trail offer a practical model for assessing AI readiness.
What Remains Unresolved
Simulated crises may not predict real-world performance, where stakes are higher and variables more complex. The benchmark also caps its scores and skips strategic planning and emotional intelligence. How configurations and training data shape these results is still unknown.
A Break from Tradition
Most AI benchmarks measure language capability or task-specific accuracy. The Crucible League is among the first to test real-world business management: social engineering resistance, trust maintenance, and decision execution under sustained stress.
Next Steps
Expect refined models with stronger document reading and follow-through, plus benchmarks featuring longer scenarios and complex decisions. Auditable decision logs are likely to become a standard feature of enterprise AI — fostering accountability in AI-driven management.
Implications for AI in Business Management
The results highlight that AI’s ability to manage complex, high-pressure business scenarios depends not only on communication skills but also on thoroughness, trustworthiness, and follow-through. While current models can recognize crises and refuse manipulation, their capacity to read and utilize internal documentation effectively remains inconsistent. This raises concerns about deploying AI in real enterprise environments where trust, integrity, and comprehensive decision-making are critical.
For organizations integrating AI agents into customer support, CRM, or decision-making, these findings underscore the importance of evaluating not just language proficiency but also operational reliability, task completion, and ethical behavior. The benchmark’s transparent scoring and auditable decision trail provide a valuable tool for assessing AI readiness in management roles.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks
Traditionally, AI performance benchmarks focus on language capabilities or task-specific accuracy, often overlooking the broader management context. The ‘Crucible League’ is among the first to evaluate AI models on their ability to handle real-world business crises, including social engineering, trust maintenance, and decision execution under stress. The benchmark was designed to simulate a week’s worth of business challenges, with full transparency and auditable decision logs to ensure fairness and accountability.
Previous assessments of AI in enterprise roles have highlighted limitations in consistency and reliability, but few have tested models in a scenario that measures their capacity to manage trust and follow-through. The July 2026 results mark a significant step toward understanding AI’s practical management potential and its current shortcomings.
It is important to note that the scores reflect not just language or superficial decision-making but also the models’ ability to read internal documentation, refuse manipulative requests, and uphold trust—key factors in real-world AI management applications.
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Performance
It is not yet clear how these benchmark results will translate to real-world enterprise environments, where variables are more complex and stakes higher. The models’ performance in simulated crises may differ from actual business management, especially regarding long-term trust and ethical decision-making. Additionally, the impact of different configurations, training data, and operational parameters on these scores remains to be studied.
Furthermore, the benchmark’s design intentionally caps scores and does not measure the full spectrum of AI management skills, such as strategic planning or emotional intelligence. Whether future models will surpass current limitations is still uncertain.
AI decision-making software for enterprises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Researchers and developers are expected to refine AI models to improve reading internal documents, follow-through, and trustworthiness. Future benchmarks may incorporate longer-term management scenarios and more complex decision-making tasks, aiming to better reflect real-world enterprise needs.
Organizations interested in deploying AI for management should monitor these developments and consider participating in or simulating similar tests to evaluate their AI tools’ readiness. The ongoing evolution of benchmarks like the Crucible League will help set industry standards for trustworthy AI management.
Additionally, transparency and auditable decision logs, as used in this benchmark, are likely to become standard features in enterprise AI solutions, fostering greater accountability and reliability in AI-driven management.
Source: ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
