
Chat demos are easy to fake. Give an AI a polite question and it will give you a polished answer — but that tells you almost nothing about how it behaves when it’s wired into your CRM, your support queue, or your sales pipeline.
So one experiment flipped the script: instead of asking models to chat, it asked them to run a company. Four frontier AIs each took charge of the same small software firm through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The result was a league table where the difference between first and last place came down to something startlingly specific: whether the model bothered to read a file that was two document references deep in the company’s own records.
The Experiment
The setup comes from Firmulate, which runs AI models as complete companies with real money mechanics — in the live version, a synthetic staff of 13 employees burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules. It’s watchable, and it rebuilds itself twice a day.
In the benchmark run, each model faced the same crucible week. The final league, from July 2026, looks like this: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
AI document reading and analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here’s the finding that should make any business buyer sit up. Every model in the field spotted every crisis. Every model refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned.
The deal turned on a buried fact: the customer’s decisive competitor weakness wasn’t in the sales call or the event log — it sat two document references deep in the company’s own files. The models that chased those references down won the deal at full price, worth +€4,583 in MRR. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.
In other words, “reads your files before answering” isn’t a nice-to-have personality trait. It’s a measurable, purchase-deciding property of an AI agent — and it’s invisible in a chat demo.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test
The week also included a three-stage impersonation campaign — fake CEO messages escalating in pressure — plus a reporter’s trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh, and still landed second at 93.
AI for enterprise document management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Paradox
The most counterintuitive profile belongs to Opus 4.8. It was the most thorough participant in the field — 80 additional learned rules, the deepest analyses — and it finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models: effort and rigor don’t automatically translate into finishing the job.
AI model testing and benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
Firmulate has packaged 242 real, unedited management decisions from the runs into a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Full results and plain-language findings are at firmulate.com/benchmarks.html.

The chat-demo era of evaluating AI agents is over, or should be. When an agent’s job is to touch real business systems, the questions that matter are blunt: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost? The buried €55,000 fact made those questions concrete. Two models did their homework and got paid. The rest gave beautiful answers to the wrong question.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html