🔍 Read the full analysis: Put AI Agents In The Hot Seat Before They Reach Your Business on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Firmulate says five AI models recognized every crisis and refused every manipulation attempt in its July 2026 Crucible League, but differed in whether they found evidence and closed a justified deal. The company offers pilots that use read-only business data to test agents without writing to live systems.
Firmulate has published results from a July 2026 wargame in which five AI models managed a simulated software company through a difficult week, reporting that all identified the crises and rejected manipulation attempts but only two signed a justified €55,000 deal. The company is also offering pilots that test models against a read-only export of a business’s data, aiming to show how agents handle company-specific pressures before they are used in live operations.
In the final Crucible League, the models made decisions for the same small software company. Firmulate says the standings were GPT-5.6-Sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The scoring allowed partial progress, while a breach of trust capped a participant’s total. Firmulate’s stated rule was: “no amount of good work outweighs a breach of trust.”
The company says each model recognized every crisis and refused every manipulation attempt, including staged fake CEO messages and a reporter’s request for an informal yes-or-no answer. The league’s largest reported gap came after the models diagnosed a sales opportunity: only two signed the deal. Firmulate says the decisive information about a competitor was buried two document references deep in company files. Models that found and used it won at full price, adding €4,583 in monthly recurring revenue.
Firmulate describes Opus 4.8 as the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The company says Opus left the deal unsigned and tried to write into a locked department rather than escalate. It reports a less severe version of that boundary issue in all four other models. Firmulate also flags a comparison caveat: Kimi K3 ran with the API’s default effort setting, while the other models ran at xhigh.
Put AI Agents in the Hot Seat Before They Reach Your Business
Firmulate’s business wargame ran five AI models through a brutal simulated week: crises, manipulation attempts, and one buried sales opportunity. All five spotted every crisis and refused every trick — but only two closed the deal. The lesson: recognition is not the same as readiness.
One League, One Simulated Company, Very Different Outcomes
All five models managed the same small software company through versioned workdays. Scoring allowed partial progress — but a breach of trust capped the total. Kimi K3 ran at the API’s default effort setting while the other models ran at xhigh, complicating direct comparison.
* Kimi K3 used the API default effort setting; other models ran at xhigh.
Seeing the Problem Is Not Solving It
Every model recognized the emergencies and resisted manipulation — including staged fake CEO messages and a reporter’s request for an informal yes-or-no answer. The decisive gap came after diagnosis: the crucial competitor information was buried two document references deep in company files. Finding it and acting on it separated the winners from the rest.
Find the buried evidence
Decisive competitor intel sat two references deep in company files. Models that located it won the deal at full price.
Close the justified deal
Same diagnosis, same pitch — no signature for three of five models, despite qualifying the opportunity.
Respect the boundaries
Opus 4.8 — the most thorough analyst, with 80 learned rules — tried to write into a locked department rather than escalate. It finished last.
Test Agents on Your Own Data — Without Touching Live Systems
Firmulate proposes running business scenarios against a read-only export of a company’s data, then delivering a board report with model rankings and playbook weaknesses. No write-back to operational systems.
Export
Company provides a read-only export of its business data.
Pressure-test
Scenarios cover customers, sales pipeline, internal rules and pressure points.
Observe
Agent decisions are recorded — nothing is written to live records.
Board report
Model rankings plus weaknesses found in existing playbooks.
What the League Actually Said
“Same diagnosis, same pitch — no signature.”
“Treat the request as a suspected approval-bypass / possible impersonation.”
“No amount of good work outweighs a breach of trust.”
What the Standings Do — and Don’t — Prove
| Question | Answer | Status |
|---|---|---|
| Crises detected | All five models recognized every crisis | ✓ |
| Manipulation refused | Fake CEO messages and reporter traps rejected | ✓ |
| Deal closed | Only two of five models signed the €55,000 deal | ~ |
| Boundary handling | All five showed some form of write-boundary issue | ✗ |
| Effort settings equal | Kimi K3 ran at default; others at xhigh | ✗ |
| General performance | One simulated company, one league — no cross-company evidence | ✗ |
| Pilot details public | Scenario list, evaluation method, and publication plans unspecified | ~ |
Quick Answers
What did the Crucible League test?
Five AI models ran the same simulated software company through a difficult week, making versioned, auditable decisions about crises, opportunities and manipulation attempts.
Which model ranked highest?
GPT-5.6-Sol with 95 points, followed by Kimi K3 with 93 — noting Kimi K3 used the default effort setting while others ran at xhigh.
What is the proposed enterprise pilot?
Scenarios run against a read-only export of a company’s data, producing a board report on model rankings and playbook weaknesses — with no write-back to live systems.
Do the results prove real-world performance?
No. Scores come from one simulated company and one league. They do not establish how models would perform across other businesses or in live operations.
From Crisis Recognition to Action
The results highlight a distinction between an agent that can identify a problem and one that can complete a suitable response. In this simulation, recognizing emergencies and resisting manipulation did not by themselves secure the deal: the outcome also depended on locating evidence already stored in company files and acting on the opportunity. That makes document retrieval, follow-through and escalation practical areas to assess before assigning agents business responsibilities.
Firmulate’s proposed pilot turns that assessment toward a company’s own information. The company says it can run scenarios against a read-only data export and produce a board report with model rankings and weaknesses in existing playbooks. Because the pilot does not write back to operational systems, it is presented as a way to examine agent behavior without letting the test make changes to live company records.
A Simulated Company Under Pressure
Firmulate’s public experiment uses a synthetic company with 13 employees, versioned workdays and playbook rules learned by the models. The company reports monthly burn of €105,000 against €2,300 in monthly recurring revenue, along with a public cash countdown. These mechanics create a constrained business setting in which decisions can be followed over time.
Readers can follow the experiment at firmulate.com and take a quiz based on 242 real, unedited management decisions, according to Firmulate. The Crucible League standings are results from this particular setup. The differing effort settings, the simulated company and the scoring rules all shape what the ranking can show; the published results do not establish how the models would perform across companies or in live deployments.
““Same diagnosis, same pitch — no signature.””
— Firmulate
Limits of the League Results
The published standings do not show whether the same models would rank similarly under different scenarios, with different company records or under live operating conditions. The effort-setting difference between Kimi K3 and the other entrants also complicates direct comparison. Firmulate’s account describes one league and its own scoring design; the results are not independent evidence of general performance across businesses.
Further details about the enterprise pilots, including which crisis scenarios are included, how company exports are prepared and how the board reports are evaluated, are not specified in the published description. It is also unclear whether pilot findings will be made public or how models will be selected for each company’s test.
Company-Specific Pilots Offered
Firmulate is inviting companies to discuss a pilot using a read-only export of their business data. The proposed exercise would test scenarios involving customers, sales pipeline, internal rules and pressure points, then provide a board report with model rankings and playbook weaknesses. The company says the export would not write back to real systems.
Readers can view the live synthetic company at firmulate.com/live and the full results at firmulate.com/benchmarks.html. Firmulate directs interested companies to its pilot page or to contact@firmulate.com. No pilot schedule, participating companies or independent evaluation results have been announced in the provided details.
Source: ThorstenMeyerAI.com
Key Questions
What did the Crucible League test?
Firmulate says five AI models ran the same simulated software company through a difficult week, making versioned and auditable decisions about crises, business opportunities and attempts to manipulate staff.
Which model ranked highest?
Firmulate’s reported standings put GPT-5.6-Sol first with 95 points, followed by Kimi K3 with 93. The company notes that Kimi K3 used the API default effort setting while the other models ran at xhigh.
What is the proposed enterprise pilot?
Firmulate says a pilot would run business scenarios against a read-only export of a company’s data and produce a board report on model rankings and weaknesses in the company’s playbooks. The company says the test does not write back to live systems.
Do the results prove how agents will perform at a company?
No. The reported scores come from one simulated company and one league. The published details do not establish how the models would perform across other businesses or in live operations.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
