firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

We live surrounded by advice about staying calm in a crisis, spotting a bad deal and knowing when to say no. But would an AI running your business do those things when the pressure is real? Firmulate turns that question into a live experiment: the same small company, the same rough week, and different AI models making the calls.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same week, different decisions

In Firmulate’s final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations. Their decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate’s standard is blunt: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models missed the danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. A polished answer, it turns out, does not guarantee a finished job.

Amazon

Top picks for "busines start company"

As an affiliate, we earn on qualifying purchases.

The clue was already in the company

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not spelled out in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In a business, useful information may be present without being obvious; an agent has to find it and follow through.

The integrity test brought a different kind of pressure. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

There were differences in how participants worked. Opus 4.8 was the most thorough, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the deal unsigned and tried to write in a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh — a useful caveat when comparing the results.

A company you can watch

The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook contains 680+ self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

That makes the experiment more than a leaderboard. It offers a running view of how AI handles competing demands: notice the crisis, respect boundaries, find the evidence and complete the task. The decisions are real within the experiment, while the employees are synthetic; the point is to observe management behavior before assigning agents work in a business.

From watching to a company-specific pilot

For an enterprise, the next step is to run the wargame against its own business context. Firmulate says a pilot can start from a read-only export, then test crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That lets leaders examine how an AI workforce might respond before it touches a CRM, support queue or forecast.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

A model can recognize a crisis and still leave the crucial decision unfinished. Firmulate’s experiment makes that gap visible, then offers enterprises a way to examine their own scenarios and procedures. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Finish What It Starts? The Crucible Experiment Reveals the Hidden Strengths of Business Bots

A live experiment reveals that AI models must do more than talk well — they need to finish what they start. Only two out of four models managed to close a critical deal under real-world stress.

Why AI Coaching Is Becoming Part of the Modern Work Stack

AI coaching is transforming workplaces with personalized, real-time guidance, but understanding its full impact requires exploring how it shapes future work environments.

What AI Coaching Gets Right About Modern Attention Spans

Clever AI coaching recognizes evolving attention spans and personalizes strategies to help you stay focused—discover how it can transform your productivity today.

Career Coaching Vs Career Counseling: Which Will Propel Your Career Forward?

Wondering whether career coaching or counseling is right for you? Discover how each can uniquely impact your career trajectory and help you decide.