firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A benchmark with customers, crises and a cash countdown

Most AI demonstrations end when the model produces a convincing answer. Firmulate asks a harder question: what happens after the answer, when software must make decisions, follow through and protect a business under pressure?

The public experiment operates a small software company staffed by 13 synthetic employees. Its financial position is deliberately uncomfortable: monthly burn stands at €105k against €2.3k in monthly recurring revenue. A public cash countdown keeps that mismatch visible. More than 680 self-learned playbook rules record what the company has learned, while every workday is versioned for inspection.

That makes Firmulate unusually watchable. It is not merely publishing a finished benchmark table; it is presenting an ongoing company story with real money mechanics and observable consequences. Readers can watch the company live as its synthetic staff confront the operational grind that polished chat demonstrations tend to omit.

The AI Sales Coach: Objection Handling, Closing, and Prospecting Reimagined (The Objection Handler's Library)

The AI Sales Coach: Objection Handling, Closing, and Prospecting Reimagined (The Objection Handler's Library)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed on the problems—and diverged on execution

For the final Crucible League in July 2026, each frontier model ran the same company through its worst week. Customers, crises and temptations remained constant. Every decision was versioned and auditable, allowing performance to be judged as management rather than conversation.

The final standings put gpt-5.6-sol at the top with 95 points, followed closely by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. For context, the do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.”

The headline result was reassuring but incomplete. All models detected every crisis, and all rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had already earned. The experiment summarizes that execution gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters because business software is rarely valuable for noticing alone. A system may accurately explain a customer’s needs, prepare a persuasive proposal and still fail to complete the action that produces revenue. In Firmulate’s test, the gap between understanding and finishing was visible in the company’s financial outcome.

The winning clue was buried in company files

The decisive competitive weakness did not appear directly in the customer event. It sat two document references deep inside the company’s own files. Models that followed those references found the fact and secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a striking finding for anyone considering AI workers around sales, support or operations. The challenge was not simply generating fluent language. Success required reading the available business record deeply enough to locate evidence that changed the negotiation. The models faced the same situation and could formulate the same broad case, but access to the crucial fact only mattered when it was discovered and used.

Pressure did not break the models’ security discipline

The experiment also subjected the models to fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result provides a useful counterweight to the failed closes. The models’ weakness was not indiscriminate compliance. They consistently protected trust under social pressure, even as some struggled with ordinary follow-through. Readers can examine more of what the synthetic employees actually said on Firmulate’s public quotes page.

Thoroughness was not enough

Opus 4.8 offers the clearest cautionary profile. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared less strongly in all four other models.

Kimi K3’s result also carries an important fairness note. It ran without an effort parameter and therefore used the API default, while the others ran at xhigh. Its 93-point finish should be read with that difference in mind.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that makes AI failure legible

Firmulate’s value is not that it turns management into a leaderboard. It makes subtle failures visible: a missed document reference, a sale that was analyzed but never signed, or a blocked action that should have been escalated. These are mundane operational details, but they determine whether an AI worker merely appears capable or produces useful work.

The broader experiment continues beyond the league table. A quiz is powered by 242 real, unedited management decisions, inviting readers to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to their real systems.

For technology watchers, the public company is the compelling part. Its 13 synthetic employees keep working, the cash countdown keeps moving and each workday adds another auditable chapter. Firmulate has turned build-in-public into something more demanding: a live test of whether autonomous software can help a fragile company survive without sacrificing trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making

MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Artificial Intelligence as an Additional Tool in Customer Relationship Management and the Impact after the COVID-19 Crisis (German Edition)

Artificial Intelligence as an Additional Tool in Customer Relationship Management and the Impact after the COVID-19 Crisis (German Edition)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

MiMo Code Launches As Open-Source Solution For AI Signal Monitoring

MiMo Code, an AI signal monitoring tool, has been released as open-source, offering operations teams a new way to track AI capability and policy shifts.

The Ghost Story Became a Forecast.

Clark’s latest essay reveals a 60% chance of automated AI research by 2028 and a 40% risk of fundamental paradigm limits, signaling major shifts ahead.

The Compute Reckoning: Anthropic Finally Admits What Customers Suspected for Ten Months

Anthropic confirmed that recent customer experience issues were due to compute shortages, after years of speculation and user frustration.

IdeaClyst: The Engine That Decides What’s Worth Building

A new idea engine called IdeaClyst analyzes roadmaps and market data to generate validated, targeted project ideas, helping founders prioritize valuable work.