firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

When evaluating AI models for critical business tasks, it’s tempting to look solely at chat scores or headline capabilities. But real-world performance often reveals unexpected insights—like the surprising fact that a do-nothing baseline AI scores 26 out of 100 in a rigorous benchmark. For business leaders, understanding why such a modest score exists—and what it truly measures—can be a game-changer in selecting and trusting AI tools.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Behind the Numbers: What a Baseline AI Score Tells Us

At first glance, a score of 26 out of 100 for a do-nothing AI may seem meaningless—after all, it’s just a baseline. But in a transparent, real-world experiment conducted by Firmulate, this number is a crucial indicator of what honest benchmarking looks like. The experiment involved running multiple advanced AI models through the simulated crisis-filled week of a small software company. Every decision was auditable, every crisis identical, and every model tested under the same conditions.

Why a do-nothing model scores 26

This baseline score is not zero because even doing nothing involves a minimal understanding of context—recognizing crises, avoiding manipulation, and maintaining discipline. It’s a reflection of the AI’s ability to identify problems without acting recklessly or making trust-breaking errors. Partial progress, like spotting a crisis but not acting or signing a deal, contributes positively; the score accumulates as the model demonstrates integrity and basic awareness.

Partial progress counts

The scoring system rewards models that can identify issues and avoid manipulation, even if they don’t fully resolve or act on them. For example, all models in the experiment identified every crisis and refused manipulation attempts, which is vital for trustworthy AI. Yet, only two models successfully signed the deal, earning full points for that achievement. The rest missed opportunities, highlighting the importance of not just spotting problems but also acting decisively.

Trust breaches cap the score

One telling fact from the experiment is that a single breach of trust—such as attempting to sign a deal they shouldn’t—capped the overall score for a model. This underscores a fundamental truth: no matter how good a model’s diagnosis or reaction, a single trust violation can undo all prior good work. This cap enforces a high standard for ethical AI behavior in critical environments.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Experiment: More Than Just Chat

Firmulate’s approach involves more than just chat demos. The models are placed in a simulated business environment with real money mechanics, a public cash countdown, and over 680 self-learned rules. The models are tested against real crises, customer manipulations, and internal document references—mirroring the complexity of actual enterprise decision-making.

In this environment, only two models, gpt-5.6-sol and Kimi K3, managed to close the €55,000 deal their own analysis had earned—meaning they trusted their diagnosis and refused to manipulate or be duped. The other models, despite reading all relevant documents, left the opportunity on the table or slipped in discipline. This highlights a key insight: the ability to read and interpret internal files, not just customer cues, can be the deciding factor in trustworthy AI performance.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust, Manipulation, and Ethical Boundaries

One of the most sophisticated tests involved social engineering—fake CEO messages escalating over multiple stages and a reporter trick asking for a quick background approval. All models refused these manipulative tactics, demonstrating an inherent understanding of trust boundaries. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

This behavior is critical for AI that interacts with human teams or customer data: it shows that models can recognize and refuse manipulation attempts, safeguarding both data and reputation.

Amazon

trustworthy AI testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Business Leaders Should Take Away

For decision-makers, the key question isn’t whether AI can produce impressive chat or write well. It’s whether AI can finish what it starts—reading critical documents, acting ethically under pressure, and maintaining trustworthiness in complex scenarios. The experiment by Firmulate demonstrates that honest benchmarks can reveal these qualities, often hidden behind shiny demos.

The current league table from this live experiment ranks models not just by raw scores but by their ability to handle crises, avoid manipulation, and close deals ethically. The top performer, gpt-5.6-sol, scored 95, followed closely by Kimi K3 at 93. The lower-ranked models still managed to close deals but slipped on discipline, scoring in the 77–88 range.

Amazon

AI benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: Why Trust and Discipline Matter

Ultimately, this transparent benchmarking approach aims to shift the focus from superficial chat capabilities to real-world trustworthiness. A do-nothing baseline scoring 26 points underscores that even minimal awareness involves ethical boundaries and careful decision-making. For businesses integrating AI, these benchmarks provide a clear-eyed view of what truly matters: honest, disciplined, and reliable AI that can navigate crises without compromising trust or integrity.

Learn more about how these models perform in real scenarios and watch the experiment unfold at firmulate.com/benchmarks.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build a Work Environment That Supports Better Decisions

Just understanding how to foster open communication, empower employees, and build trust can transform your workplace into a hub of smarter decision-making.

Why AI Coaching Feels Different When It Learns Your Patterns

Why does AI coaching feel more personal when it learns your patterns? Discover how tailored support transforms your growth journey.

Unlock Your Full Potential with the AI Coach Everyone’s Talking About

Elevate your performance and growth with the AI Coach that’s on everyone’s lips. Begin your journey to full potential today!

How AI Coaching Can Improve Follow-Through More Than Inspiration

Just understanding how AI coaching enhances follow-through can transform your approach—discover the key differences that make lasting change possible.