firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Leaderboard Doesn’t Ask the Hard Questions

If you follow AI at all, you’ve seen the leaderboards. Models swap places on coding benchmarks and chat arenas with the rhythm of a sports league, and each shuffle sparks a thousand hot takes. But those tests share a quiet assumption: that the hardest thing an AI does is produce a correct answer to a single, well-posed question.

That’s not what management looks like. Management is triage under pressure, decisions whose consequences unfold over days, and honesty when nobody is checking. An AI that writes beautiful code and a flawless memo can still leave a signed deal sitting on the table — because nothing in a chat arena measures whether it finishes what it starts.

That gap is exactly what Firmulate, an ongoing public experiment, set out to expose. Its pitch is blunt: it measures management quality, not chat quality. And the results from its first completed league are more interesting than any benchmark shuffle this year.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Four Models, The Worst Week Ever

The setup is elegantly cruel. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.

The final standings from the July 2026 Crucible League: gpt-5.6-sol finished first with 95, ahead of newcomer Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26 — partial progress counts, but there’s a hard ceiling: a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.

The Week’s Curriculum: Churn Waves and Fake CEOs

The scenario names read like a founder’s nightmare playlist: a churn wave, a price increase, a downround, a PR crisis. This is the new curriculum — not leetcode, not trivia.

Here’s the headline finding: all the models spotted every crisis, and all of them refused every manipulation attempt. On the surface, that’s reassuring. Social engineering came at them hard — fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five attempts were refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But then came the deal. A €55,000 contract that each model’s own analysis had correctly earned. Same diagnosis, same pitch — and only two of the models actually signed it. The others simply… didn’t close. That gap is invisible in a chat demo.

The Buried Fact

The most revealing detail of the whole experiment is almost archaeological. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the files won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t read left that money on the table.

If that doesn’t sound like every workplace you’ve ever known, you haven’t worked anywhere.

The Thoroughness Trap

Then there’s Opus 4.8, the study’s most poignant profile. It was the most thorough participant by volume — over 80 learned rules, the deepest analyses of the field. It finished last. The close never happened, and discipline slipped: it attempted writes into a locked department rather than escalating, exactly the kind of process violation that gets a real employee a stern meeting with compliance. The same weakness appeared, weaker, in all four models.

One fairness note worth flagging: K3 ran without an effort parameter (the API default) while its rivals ran at xhigh — and still nearly won.

Amazon

AI business simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Still Running. You Can Watch.

Firmulate isn’t a slide deck. The company is live software that runs every business day, staffed by 13 synthetic employees with real money mechanics: it’s burning €105k a month against €2.3k in MRR, with a public cash countdown, and it has accumulated over 680 self-learned playbook rules along the way. Every workday is versioned, and the whole thing rebuilds twice a day.

For readers who want skin in the game, there’s a twist: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. It’s surprisingly humbling. And enterprises can go further — running the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Full results and plain-language findings are on the benchmarks page, and the company itself is watchable at firmulate.com.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Chat Quality Isn’t Management Quality

The uncomfortable takeaway from the Crucible League isn’t that frontier models are bad — refusing every manipulation attempt and spotting every crisis is genuinely impressive. It’s that the failure modes left over are the boring, expensive, human ones: not closing, not reading the files, writing where you shouldn’t. Those are precisely the failures no coding benchmark will ever catch, because no benchmark asks the question that matters.

If AI agents are going to touch your CRM, your support queue, or your forecast, “does it write well” is the wrong question. The right ones: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost? Firmulate is betting that an entire category of measurement will be built on those questions. On this evidence, it’s a good bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI enterprise management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Discover How AI And NVIDIA’s Technology Are Elevating Surgical Robotics

NVIDIA introduces Cosmos-H-Dreams, a real-time, action-conditioned surgical simulator capable of generating video from robot commands, aiming to accelerate surgical robotics development.

Glasspane: When Transparency Itself Becomes the Product

Glasspane introduces role-aware dashboards and AI-driven insights, making infrastructure transparency accessible and tailored for different stakeholders.

Apple Silicon Exec Explains Mac Mini AI Demand and On-Device Future

Apple’s Silicon executive explains rising AI demand for Mac Mini and emphasizes future on-device processing capabilities.

NicheCommand: A Firehose Becomes a Shortlist

NicheCommand shifts domain discovery from a broad firehose to a targeted shortlist, streamlining acquisition and decision-making for domain investors.