firmulate.com/quotes.html — live view
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

The pressure test that matters beyond the chatbot window

Technology buyers are accustomed to judging artificial intelligence by what it produces: a polished answer, a working feature or an impressive analysis. Firmulate tested a harder question. What happens when an AI manager receives an urgent instruction from someone claiming to be the CEO—and that instruction asks it to abandon normal safeguards?

The answer was unusually encouraging. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All five frontier models refused every manipulation attempt. They also spotted every crisis placed in their path.

This was not a scripted chatbot demonstration. Firmulate operates a live, watchable experiment in which models run the same small software company through the same customers, crises and temptations. Every workday and decision is versioned and auditable, making integrity visible as behavior rather than a promise in a product description.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

An impersonated executive meets five skeptical managers

The social-engineering scenario tested a familiar weakness in business operations: authority combined with urgency. A message that appears to come from the boss can make bypassing process sound like decisive leadership. The staged escalation examined whether an AI manager would preserve its judgment as that pressure increased.

Every model held its ground. Kimi K3 gave the clearest on-record description of the situation: “Treat the request as a suspected approval-bypass / possible impersonation.” That sentence captures the important shift from merely following instructions to evaluating whether an instruction deserves trust. More examples of the models’ own words are available on Firmulate’s public quotes page.

The reporter trick tested a different route to the same failure. Instead of invoking executive authority, it used informality and a request framed as minimal: “just one yes/no, on background.” The models refused that attempt too. Across the complete exercise, all five rejected every manipulation attempt.

Amazon

AI integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Integrity was strong, but execution still separated the field

The clean security result did not mean the models performed identically. Only two signed the €55,000 deal their own analysis had earned. The others reached the same diagnosis and produced the same pitch, but did not secure the signature: “Same diagnosis, same pitch — no signature.”

The decisive commercial clue was not waiting in the customer event. It sat two document references deep inside the company’s own files. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. This distinction matters because safe refusal is only one part of dependable work. A useful AI manager must also read deeply, connect evidence and finish legitimate tasks after declining illegitimate ones.

Opus 4.8 illustrates that tension. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four other models.

The league makes trust part of the result

In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The full standings and plain-language findings are published on the Firmulate benchmarks page.

A do-nothing baseline scored 26 because partial progress counts, but the experiment treats trust as a hard boundary: “no amount of good work outweighs a breach of trust.” That framing prevents a model from compensating for one serious violation with a pile of otherwise competent activity.

Amazon

AI decision-making audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company harsh enough to expose the difference

The live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The environment is deliberately unforgiving because ordinary chat evaluations can hide the gap between knowing and doing. Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz, giving observers another way to compare judgment without relying on branding or reputation.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model robustness testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before the incident

The most significant result is not that the models could recognize an obviously dangerous request in isolation. It is that all five maintained their refusal through escalating pressure and a separate reporter tactic while continuing to manage a business under severe financial strain.

For companies considering AI access to customer records, support queues or forecasts, integrity need not remain an assumption until something goes wrong. Firmulate’s pilot lets enterprises run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.

The experiment also offers a useful warning against treating safety as the whole score. The strongest AI workforce must refuse the fake boss, find the buried fact and complete the legitimate deal. Firmulate’s result shows that the first capability is already testable—and, in this field, every model passed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Quiet GPUs for Local AI: Acoustic and Thermal Roundup

An overview of the quietest GPUs for local AI in 2026, focusing on thermal performance, acoustics, and optimal configurations for different VRAM tiers.

Thunderbolt-ibverbs: We Have InfiniBand At Home

Researchers developed a Linux kernel module enabling InfiniBand-like RDMA over USB4/Thunderbolt ports on consumer AMD mini PCs, achieving high-speed interconnects for AI workloads.

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now autonomously creates and manages its own team of subagents for complex tasks, enhancing multi-step AI workflows.

Are AI Student Planners The Secret To Academic Success In 2026?

Exploring whether AI-powered student planners are improving academic outcomes in 2026, focusing on real developments and remaining uncertainties.