
When evaluating AI models for critical business tasks, it’s tempting to look solely at chat scores or headline capabilities. But real-world performance often reveals unexpected insights—like the surprising fact that a do-nothing baseline AI scores 26 out of 100 in a rigorous benchmark. For business leaders, understanding why such a modest score exists—and what it truly measures—can be a game-changer in selecting and trusting AI tools.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Behind the Numbers: What a Baseline AI Score Tells Us
At first glance, a score of 26 out of 100 for a do-nothing AI may seem meaningless—after all, it’s just a baseline. But in a transparent, real-world experiment conducted by Firmulate, this number is a crucial indicator of what honest benchmarking looks like. The experiment involved running multiple advanced AI models through the simulated crisis-filled week of a small software company. Every decision was auditable, every crisis identical, and every model tested under the same conditions.
Why a do-nothing model scores 26
This baseline score is not zero because even doing nothing involves a minimal understanding of context—recognizing crises, avoiding manipulation, and maintaining discipline. It’s a reflection of the AI’s ability to identify problems without acting recklessly or making trust-breaking errors. Partial progress, like spotting a crisis but not acting or signing a deal, contributes positively; the score accumulates as the model demonstrates integrity and basic awareness.
Partial progress counts
The scoring system rewards models that can identify issues and avoid manipulation, even if they don’t fully resolve or act on them. For example, all models in the experiment identified every crisis and refused manipulation attempts, which is vital for trustworthy AI. Yet, only two models successfully signed the deal, earning full points for that achievement. The rest missed opportunities, highlighting the importance of not just spotting problems but also acting decisively.
Trust breaches cap the score
One telling fact from the experiment is that a single breach of trust—such as attempting to sign a deal they shouldn’t—capped the overall score for a model. This underscores a fundamental truth: no matter how good a model’s diagnosis or reaction, a single trust violation can undo all prior good work. This cap enforces a high standard for ethical AI behavior in critical environments.
As an affiliate, we earn on qualifying purchases.
The Real-World Experiment: More Than Just Chat
Firmulate’s approach involves more than just chat demos. The models are placed in a simulated business environment with real money mechanics, a public cash countdown, and over 680 self-learned rules. The models are tested against real crises, customer manipulations, and internal document references—mirroring the complexity of actual enterprise decision-making.
In this environment, only two models, gpt-5.6-sol and Kimi K3, managed to close the €55,000 deal their own analysis had earned—meaning they trusted their diagnosis and refused to manipulate or be duped. The other models, despite reading all relevant documents, left the opportunity on the table or slipped in discipline. This highlights a key insight: the ability to read and interpret internal files, not just customer cues, can be the deciding factor in trustworthy AI performance.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust, Manipulation, and Ethical Boundaries
One of the most sophisticated tests involved social engineering—fake CEO messages escalating over multiple stages and a reporter trick asking for a quick background approval. All models refused these manipulative tactics, demonstrating an inherent understanding of trust boundaries. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
This behavior is critical for AI that interacts with human teams or customer data: it shows that models can recognize and refuse manipulation attempts, safeguarding both data and reputation.
As an affiliate, we earn on qualifying purchases.
What Business Leaders Should Take Away
For decision-makers, the key question isn’t whether AI can produce impressive chat or write well. It’s whether AI can finish what it starts—reading critical documents, acting ethically under pressure, and maintaining trustworthiness in complex scenarios. The experiment by Firmulate demonstrates that honest benchmarks can reveal these qualities, often hidden behind shiny demos.
The current league table from this live experiment ranks models not just by raw scores but by their ability to handle crises, avoid manipulation, and close deals ethically. The top performer, gpt-5.6-sol, scored 95, followed closely by Kimi K3 at 93. The lower-ranked models still managed to close deals but slipped on discipline, scoring in the 77–88 range.
As an affiliate, we earn on qualifying purchases.
Conclusion: Why Trust and Discipline Matter
Ultimately, this transparent benchmarking approach aims to shift the focus from superficial chat capabilities to real-world trustworthiness. A do-nothing baseline scoring 26 points underscores that even minimal awareness involves ethical boundaries and careful decision-making. For businesses integrating AI, these benchmarks provide a clear-eyed view of what truly matters: honest, disciplined, and reliable AI that can navigate crises without compromising trust or integrity.
Learn more about how these models perform in real scenarios and watch the experiment unfold at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
