firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A convincing demo can hide the decision that matters

An AI assistant can spot a crisis and deliver a polished pitch. But when the moment comes to act, will it follow through—and respect the rules while it does? Firmulate puts that question to work in a live experiment: AI models run the same small software company through a bad week, with real money mechanics and decisions people can watch.

Same company, same pressure

In the final Crucible League, completed in July 2026, frontier models faced the same customers, crises and temptations. Every decision was versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The standard was deliberately unforgiving: partial progress counted, but one breach of trust capped the total. “No amount of good work outweighs a breach of trust.”

On crisis recognition and resisting manipulation, the models all cleared the bar: every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The summary is striking: “Same diagnosis, same pitch — no signature.” A chat demonstration might show the diagnosis and the pitch. The live company experiment exposes whether the model completes the consequential step.

The clue was already in the files

The deciding weakness in a competitor’s position was buried two document references deep in the company’s own files. It did not appear in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result turns a familiar promise about AI agents into a practical question for business leaders: can a system find and use relevant information already available to it when a real decision is on the line?

The experiment also staged social-engineering pressure. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong analysis is not the whole job

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped when it attempted writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four participants. K3 also ran without an effort parameter, using the API default; the others ran at xhigh. That difference is relevant context when comparing the results.

The live company makes the stakes visible in a different way. It has 13 synthetic employees, a public cash countdown, 680+ self-learned playbook rules and versioned work every business day. Its finances show burn of €105k/month against €2.3k MRR. Firmulate describes the experiment as real and watchable at firmulate.com. A separate quiz uses 242 real, unedited management decisions to let readers guess which model made them.

From watching to trying it on your business

For an enterprise, the next step is a pilot built from a read-only export of its own business. Teams can put crisis scenarios against a digital twin of their company and get a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That offers a way to examine how AI might handle a company’s pressures before trusting it with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your own playbooks hold up

Firmulate’s live experiment shows why spotting a crisis is only part of the job: an AI system also has to act on what it finds, close the loop and keep its discipline. Enterprises can run the wargame against a read-only export of their business. Explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Real Cost Of A Local-Inference Rig In 2026

Analyzing the costs and hardware requirements for local AI inference in 2026, highlighting VRAM constraints, hardware choices, and value metrics.

Exploring The Top 8 AI Drawing Tablets In 2026

Discover the leading AI drawing tablets of 2026, featuring top models for beginners and professionals, with detailed insights on features and performance.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX’s $60 billion all-stock purchase of AI coding startup Cursor is a strategic move, offering growth and competitive advantages amid rapid revenue growth.

Your Company Data And AI In 2026: Insights Into OpenAI’s Data Framework

OpenAI clarifies its data handling policies in 2026, emphasizing no default training on enterprise data and new governance controls for AI agents.