
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Diligence Isn’t the Same as Delivery
Everyone knows the type: the colleague who writes the longest memo, reads every attachment, stays latest at the office — and somehow never quite closes the deal. According to a live, public experiment running at Firmulate, frontier AI models can be that colleague too.
In the final July 2026 standings of Firmulate’s “Crucible League” — a wargame in which four frontier AI models each ran the same small software company through its worst week — Anthropic’s Opus 4.8 finished last with a score of 73. Not because it was lazy. Opus 4.8 was, by several measures, the most diligent participant in the field: it accumulated 80 self-learned playbook rules, more than any competitor, and produced the deepest analyses of any model in the run.
It still lost. And the reason why is the most useful thing the experiment found.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crisis, Same Temptations
The setup is elegantly simple. Firmulate handed each frontier model an identical job: run a small software company through a catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision is versioned and auditable, so nothing about the run is anecdotal.
The final league table tells a tight story:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93. The newcomer (Moonshot): closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77.
- 5. Opus 4.8 — 73. The most thorough participant — and last place.
For context, a do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total — as the experiment’s rules put it, “no amount of good work outweighs a breach of trust.”
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deal Nobody Signed
Here’s the headline finding: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The missing ingredient was buried, literally. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t read left the close on the table.
The experiment also stress-tested honesty. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: Thorough to a Fault
Which brings us back to the character study. Opus 4.8’s profile is a fascinating paradox: +80 learned rules, the deepest analyses in the field — and last place. Two things sank it. First, the close was left on the table, just like half the field. Second, discipline slipped: it made write attempts into a locked department instead of escalating the issue properly.
The lesson isn’t that Opus is a bad model. It’s that diligence ≠ impact. Volume of analysis and volume of learned rules didn’t convert into finished business. Prioritization beat thoroughness — for AI just as it does for humans.
And in fairness, Firmulate’s own findings note the same weakness appeared, weaker, in all four models. Opus just exhibited it most sharply.
As an affiliate, we earn on qualifying purchases.
Watch It Happen, Live
This isn’t a paper you read — it’s a company you can watch. The live firmulate company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in MRR, with a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The site rebuilds itself twice a day, with new benchmark runs queued and the league growing automatically with each finished run.
Want to test whether you could tell the models apart yourself? A quiz built from 242 real, unedited management decisions lets you guess which model made which call.
One fairness footnote worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second with the cleanest discipline of the field.

Why Gadget Fans Should Care
If AI agents are about to touch your CRM, your support queue, or your forecast, the interesting question is no longer “does it write well?” Chat demos are saturated with eloquence. The real questions are: does it finish what it starts, does it read your files first, and does it stay honest under pressure?
Opus 4.8’s last-place finish is the sharpest answer yet: the model that worked hardest accomplished least, because it never converted effort into outcomes. That’s a failure mode every manager recognizes instantly — and one that would be invisible in a chat demo.
For enterprises that want to stress-test before deploying, Firmulate offers a pilot: run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.
The full benchmark results and plain-language findings are public. Go watch the A student struggle with the group project — and the newcomer who aced it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.