firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Chat demos are easy to fake. Give an AI a polite question and it will give you a polished answer — but that tells you almost nothing about how it behaves when it’s wired into your CRM, your support queue, or your sales pipeline.

So one experiment flipped the script: instead of asking models to chat, it asked them to run a company. Four frontier AIs each took charge of the same small software firm through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The result was a league table where the difference between first and last place came down to something startlingly specific: whether the model bothered to read a file that was two document references deep in the company’s own records.

The Experiment

The setup comes from Firmulate, which runs AI models as complete companies with real money mechanics — in the live version, a synthetic staff of 13 employees burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules. It’s watchable, and it rebuilds itself twice a day.

In the benchmark run, each model faced the same crucible week. The final league, from July 2026, looks like this: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

Here’s the finding that should make any business buyer sit up. Every model in the field spotted every crisis. Every model refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned.

The deal turned on a buried fact: the customer’s decisive competitor weakness wasn’t in the sales call or the event log — it sat two document references deep in the company’s own files. The models that chased those references down won the deal at full price, worth +€4,583 in MRR. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.

In other words, “reads your files before answering” isn’t a nice-to-have personality trait. It’s a measurable, purchase-deciding property of an AI agent — and it’s invisible in a chat demo.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Social Engineering Test

The week also included a three-stage impersonation campaign — fake CEO messages escalating in pressure — plus a reporter’s trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh, and still landed second at 93.

Amazon

AI for enterprise document management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

The most counterintuitive profile belongs to Opus 4.8. It was the most thorough participant in the field — 80 additional learned rules, the deepest analyses — and it finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models: effort and rigor don’t automatically translate into finishing the job.

Amazon

AI model testing and benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try It Yourself

Firmulate has packaged 242 real, unedited management decisions from the runs into a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The chat-demo era of evaluating AI agents is over, or should be. When an agent’s job is to touch real business systems, the questions that matter are blunt: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost? The buried €55,000 fact made those questions concrete. Two models did their homework and got paid. The rest gave beautiful answers to the wrong question.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Home signal monitor: Mortgage Rates Inch to Another 6-Week Low

Mortgage rates have decreased to their lowest level in six weeks, potentially impacting homebuyers and the housing market.

Parent-teacher Meeting Prep Brief

A new workflow for elementary teachers to streamline parent meeting prep is being tested, aiming to reduce planning time with a simple digital tool.

Sovereignty Is a Pipe, Not a Passport

Exploring how data sovereignty depends on legal jurisdiction, not physical location, with implications for European AI providers like Mistral.

A Skill Is a Folder, Not a Prompt: What Anthropic Learned Running Hundreds of Them

Anthropic reveals that effective AI skills are structured as folders containing instructions, scripts, and assets—transforming prompt engineering into durable organizational assets.