firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

Frontier AI models make surprisingly recognizable managers

Technology fans have become adept at spotting the fingerprints of different gadgets, operating systems and chatbots. Firmulate poses a harder question: can you identify an AI model from a consequential management decision rather than its writing style?

Its guess-the-model quiz draws on 242 real, unedited decisions made during a live business experiment. Each model confronted the same small software company, customers, crises and temptations during its worst week. Readers see what the managers actually decided, then guess which frontier model was responsible.

The result feels part personality test, part management case study. One model may produce exhaustive analysis, another may respond with striking brevity, and another may reject a seemingly harmless request because it detects an attempt to bypass proper approval. These are not fictional characters written for entertainment. They are observable differences in how AI systems behave when placed in identical situations.

Amazon

AI management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business wargame, not a chat demo

Firmulate runs a synthetic company with 13 employees and real money mechanics. The business burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The central test was deliberately unforgiving. Every frontier model had to manage the same emergencies while preserving customer trust and resisting manipulation. The do-nothing baseline scored 26 because partial progress counts, but one trust breach caps the total. The experiment’s standard is blunt: “no amount of good work outweighs a breach of trust.”

In the final Crucible League results from July 2026, gpt-5.6-sol led with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. The ranking makes the quiz more than a game of matching prose styles. It offers a window into the behaviors behind materially different outcomes.

The difference between seeing work and finishing it

All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap as “Same diagnosis, same pitch — no signature.”

That failure is easy to miss in a conventional chatbot demonstration. A fluent answer can sound complete even when the necessary business action remains undone. In Firmulate’s company, recognizing the opportunity, preparing the argument and actually closing the deal were separate tests of management quality.

The decisive information was also hidden in a revealing place. A competitor weakness sat two document references deep in the company’s own files rather than in the customer event. Models that found and used that fact won the contract at full price, adding €4,583 in monthly recurring revenue. The finding turns a familiar workplace instruction—read the files—into a measurable competitive advantage.

Pressure revealed discipline as well as competence

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest security-minded explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because the requests were designed to exploit authority, urgency and informality rather than technical weakness. The models did not merely recognize operational problems; they maintained boundaries when someone appeared to offer an easy shortcut.

Kimi K3’s strong finish deserves one qualification. It ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase its result, but it is important context when comparing behavior across the field.

Why the most diligent model finished last

Opus 4.8 provides the experiment’s most interesting cautionary tale. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses. It nevertheless finished last because it left the deal unsigned and lost process discipline.

One recurring problem was attempting to write into a locked department instead of escalating the blockage. A weaker version of that behavior appeared in all four of the other models. Thoroughness, in other words, did not guarantee completion—and repeatedly trying a blocked route was not a substitute for managerial judgment.

Infographic —
The findings at a glance — source: firmulate.com.
AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz tests the traits that benchmarks often hide

Firmulate’s experiment suggests that AI management personalities can be observed through follow-through, reading habits, escalation choices and resistance to social pressure. Those differences become especially visible when every participant receives the same situation and its decisions remain auditable.

For readers, the quiz offers an unusually accessible way to inspect those distinctions firsthand. For organizations considering AI agents, the broader lesson is practical: evaluate what a model does across a complete workflow, not merely how convincing its first answer sounds.

Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing teams to examine how an AI workforce might behave before granting it operational authority. The live company keeps the underlying experiment watchable; the quiz turns its management record into a challenge anyone can try.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation games

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR Benchmark: There Is No Best Model

VigilSAR Benchmark reveals there is no universally best AI model for defense, emphasizing context-specific rankings based on capability, reliability, and compliance.

The Machine Economy — Capital-Heavy, Human-Light, Trading With Itself

Analysis of how AI-driven firms are evolving into autonomous, capital-intensive entities, reshaping economic structures and raising new policy challenges.

The calendar technicality. Why Elon Musk’s lawsuit against Sam Altman and OpenAI lost on timing, not on substance.

Elon Musk’s lawsuit claiming OpenAI violated charitable trust law was dismissed on procedural grounds, not on the case’s merits, opening IPO prospects.

The Bold Move: Launching AI-Integrated Health Features In ChatGPT

OpenAI introduces ChatGPT Health, enabling users to connect medical records and wellness apps for personalized health responses, with privacy safeguards.