firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Diligence Isn’t the Same as Delivery

Everyone knows the type: the colleague who writes the longest memo, reads every attachment, stays latest at the office — and somehow never quite closes the deal. According to a live, public experiment running at Firmulate, frontier AI models can be that colleague too.

In the final July 2026 standings of Firmulate’s “Crucible League” — a wargame in which four frontier AI models each ran the same small software company through its worst week — Anthropic’s Opus 4.8 finished last with a score of 73. Not because it was lazy. Opus 4.8 was, by several measures, the most diligent participant in the field: it accumulated 80 self-learned playbook rules, more than any competitor, and produced the deepest analyses of any model in the run.

It still lost. And the reason why is the most useful thing the experiment found.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crisis, Same Temptations

The setup is elegantly simple. Firmulate handed each frontier model an identical job: run a small software company through a catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision is versioned and auditable, so nothing about the run is anecdotal.

The final league table tells a tight story:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93. The newcomer (Moonshot): closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant — and last place.

For context, a do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total — as the experiment’s rules put it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal Nobody Signed

Here’s the headline finding: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The missing ingredient was buried, literally. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t read left the close on the table.

The experiment also stress-tested honesty. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: Thorough to a Fault

Which brings us back to the character study. Opus 4.8’s profile is a fascinating paradox: +80 learned rules, the deepest analyses in the field — and last place. Two things sank it. First, the close was left on the table, just like half the field. Second, discipline slipped: it made write attempts into a locked department instead of escalating the issue properly.

The lesson isn’t that Opus is a bad model. It’s that diligence ≠ impact. Volume of analysis and volume of learned rules didn’t convert into finished business. Prioritization beat thoroughness — for AI just as it does for humans.

And in fairness, Firmulate’s own findings note the same weakness appeared, weaker, in all four models. Opus just exhibited it most sharply.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch It Happen, Live

This isn’t a paper you read — it’s a company you can watch. The live firmulate company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in MRR, with a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The site rebuilds itself twice a day, with new benchmark runs queued and the league growing automatically with each finished run.

Want to test whether you could tell the models apart yourself? A quiz built from 242 real, unedited management decisions lets you guess which model made which call.

One fairness footnote worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second with the cleanest discipline of the field.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why Gadget Fans Should Care

If AI agents are about to touch your CRM, your support queue, or your forecast, the interesting question is no longer “does it write well?” Chat demos are saturated with eloquence. The real questions are: does it finish what it starts, does it read your files first, and does it stay honest under pressure?

Opus 4.8’s last-place finish is the sharpest answer yet: the model that worked hardest accomplished least, because it never converted effort into outcomes. That’s a failure mode every manager recognizes instantly — and one that would be invisible in a chat demo.

For enterprises that want to stress-test before deploying, Firmulate offers a pilot: run the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

The full benchmark results and plain-language findings are public. Go watch the A student struggle with the group project — and the newcomer who aced it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Shift will clean homes for free to train future robots

Shift provides free home cleaning services in exchange for recording cleaning tasks to train AI robots, raising privacy and ethical questions.

Is AI causing a repeat of Front end’s Lost Decade?

Analysis of how AI’s impact on programming mirrors past front-end deskilling, and what this means for developers and the industry.

The Compounding Error Problem — Why 99.9% Alignment Decays to 60% in 500 Generations

Research indicates that even 99.9% alignment accuracy per generation drops to around 60% after 500 recursive AI generations, raising concerns about long-term safety.

Analyzing AI Compression Workflows: Focus On Local LLMs In 2026

Exploring how native quantization-aware training shapes local large language model deployment in 2026, with focus on new low-precision formats and workflows.