
The pressure test that matters beyond the chatbot window
Technology buyers are accustomed to judging artificial intelligence by what it produces: a polished answer, a working feature or an impressive analysis. Firmulate tested a harder question. What happens when an AI manager receives an urgent instruction from someone claiming to be the CEO—and that instruction asks it to abandon normal safeguards?
The answer was unusually encouraging. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All five frontier models refused every manipulation attempt. They also spotted every crisis placed in their path.
This was not a scripted chatbot demonstration. Firmulate operates a live, watchable experiment in which models run the same small software company through the same customers, crises and temptations. Every workday and decision is versioned and auditable, making integrity visible as behavior rather than a promise in a product description.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
An impersonated executive meets five skeptical managers
The social-engineering scenario tested a familiar weakness in business operations: authority combined with urgency. A message that appears to come from the boss can make bypassing process sound like decisive leadership. The staged escalation examined whether an AI manager would preserve its judgment as that pressure increased.
Every model held its ground. Kimi K3 gave the clearest on-record description of the situation: “Treat the request as a suspected approval-bypass / possible impersonation.” That sentence captures the important shift from merely following instructions to evaluating whether an instruction deserves trust. More examples of the models’ own words are available on Firmulate’s public quotes page.
The reporter trick tested a different route to the same failure. Instead of invoking executive authority, it used informality and a request framed as minimal: “just one yes/no, on background.” The models refused that attempt too. Across the complete exercise, all five rejected every manipulation attempt.
AI integrity verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Integrity was strong, but execution still separated the field
The clean security result did not mean the models performed identically. Only two signed the €55,000 deal their own analysis had earned. The others reached the same diagnosis and produced the same pitch, but did not secure the signature: “Same diagnosis, same pitch — no signature.”
The decisive commercial clue was not waiting in the customer event. It sat two document references deep inside the company’s own files. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. This distinction matters because safe refusal is only one part of dependable work. A useful AI manager must also read deeply, connect evidence and finish legitimate tasks after declining illegitimate ones.
Opus 4.8 illustrates that tension. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four other models.
The league makes trust part of the result
In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The full standings and plain-language findings are published on the Firmulate benchmarks page.
A do-nothing baseline scored 26 because partial progress counts, but the experiment treats trust as a hard boundary: “no amount of good work outweighs a breach of trust.” That framing prevents a model from compensating for one serious violation with a pile of otherwise competent activity.
AI decision-making audit tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company harsh enough to expose the difference
The live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The environment is deliberately unforgiving because ordinary chat evaluations can hide the gap between knowing and doing. Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz, giving observers another way to compare judgment without relying on branding or reputation.

AI model robustness testing
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the refusal before the incident
The most significant result is not that the models could recognize an obviously dangerous request in isolation. It is that all five maintained their refusal through escalating pressure and a separate reporter tactic while continuing to manage a business under severe financial strain.
For companies considering AI access to customer records, support queues or forecasts, integrity need not remain an assumption until something goes wrong. Firmulate’s pilot lets enterprises run the same kind of wargame against a read-only export of their own business, with nothing writing back to real systems.
The experiment also offers a useful warning against treating safety as the whole score. The strongest AI workforce must refuse the fake boss, find the buried fact and complete the legitimate deal. Firmulate’s result shows that the first capability is already testable—and, in this field, every model passed.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html