
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Model Nobody Was Watching Nearly Won the Whole Thing
If you’ve been tracking the AI race as a two-horse game between OpenAI and Anthropic, July just got uncomfortable. In Firmulate‘s Crucible league — a live experiment where frontier AI models each run the same small software company through its worst week — Moonshot’s Kimi K3 finished second with a score of 93, beating Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol edged it out, at 95.
And K3 did it with what the experiment’s benchmark calls the cleanest discipline in the field: exactly one deviation across the entire week.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crises, Same Temptations
The setup is simple and brutal. Each frontier model was handed the same small software company — 13 synthetic employees, real money mechanics, a burn rate of €105k/month against just €2.3k in monthly recurring revenue, and a public cash countdown. Same customers, same crises, same temptations to cut corners. Only the model changes, and every decision is versioned and auditable. You can watch it play out at firmulate.com/live.
The headline finding wasn’t about raw intelligence. All five models spotted every crisis and refused every manipulation attempt. The gap showed up at the finish line: only two models signed the €55,000 deal their own analysis had earned. As the experiment’s own summary puts it: “Same diagnosis, same pitch — no signature.”
AI decision-making simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Needle
What separated the winners from the also-rans? A single decisive fact, deliberately buried two document references deep in the company’s own files — a competitor weakness that had nothing to do with the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in MRR. The models that didn’t, didn’t.
K3 found it. It also saved the churning customer and resisted all three social-engineering baits: fake CEO messages escalating over three stages, plus a reporter offering the classic “just one yes/no, on background” trick. All five models refused that one — but K3 left an on-record reasoning note worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”
For context, the do-nothing baseline scores 26, and partial progress counts — but a single breach of trust caps the total. As the league’s rule puts it, “no amount of good work outweighs a breach of trust.”
AI ethics and trustworthiness assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Harder You Try, the Worse You Do?
The most awkward result belongs to Opus 4.8: the most thorough participant in the field, with 80-plus newly learned rules and the deepest analyses — and a last-place score of 73. It left the close on the table, and its discipline slipped, including write attempts into a locked department instead of escalating the issue. Notably, the same weakness appeared, in weaker form, across all four other models.
One fairness caveat, in the interest of transparency: K3 ran without an effort parameter (API default), while the other four models ran at xhigh. In other words, the newcomer may not even have been trying its hardest.
AI enterprise decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
The experiment is unusually hands-on for the public. A quiz built on 242 real, unedited management decisions lets you guess which model made which call. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full league tables and plain-language findings are on the benchmarks page.

Why This Matters
The Crucible league makes one thing plain: the frontier is no longer a Western club. A newcomer from Moonshot walked into the hardest management simulation available and beat three of four established giants — while possibly running below its maximum effort setting.
The broader lesson is about how you choose a model at all. Chat demos measure eloquence. This experiment measures whether an AI finishes what it starts, reads your files before acting, and stays honest under pressure. Those are exactly the traits that matter if an agent will touch your CRM, support queue, or forecast — and they’re invisible in a conversation window.
If the gap between first and fifth place is 22 points on decisions that any vendor’s demo would look fine making, then picking a model without testing it against your own business is no longer a procurement choice. It’s a bet. Firmulate just made that bet checkable.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
