firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

We live surrounded by advice about staying calm in a crisis, spotting a bad deal and knowing when to say no. But would an AI running your business do those things when the pressure is real? Firmulate turns that question into a live experiment: the same small company, the same rough week, and different AI models making the calls.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same week, different decisions

In Firmulate’s final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations. Their decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate’s standard is blunt: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models missed the danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. A polished answer, it turns out, does not guarantee a finished job.

Amazon

Top picks for "busines start company"

As an affiliate, we earn on qualifying purchases.

The clue was already in the company

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not spelled out in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In a business, useful information may be present without being obvious; an agent has to find it and follow through.

The integrity test brought a different kind of pressure. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

There were differences in how participants worked. Opus 4.8 was the most thorough, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the deal unsigned and tried to write in a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh — a useful caveat when comparing the results.

A company you can watch

The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook contains 680+ self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

That makes the experiment more than a leaderboard. It offers a running view of how AI handles competing demands: notice the crisis, respect boundaries, find the evidence and complete the task. The decisions are real within the experiment, while the employees are synthetic; the point is to observe management behavior before assigning agents work in a business.

From watching to a company-specific pilot

For an enterprise, the next step is to run the wargame against its own business context. Firmulate says a pilot can start from a read-only export, then test crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That lets leaders examine how an AI workforce might respond before it touches a CRM, support queue or forecast.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

A model can recognize a crisis and still leave the crucial decision unfinished. Firmulate’s experiment makes that gap visible, then offers enterprises a way to examine their own scenarios and procedures. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build a Creator Workspace That Supports Real Consistency

AIThis post was created with the assistance of artificial intelligence (AI).To build…

Real‑Time Sentiment Analysis: How Your AI Coach Senses Your Mood

Learning how your AI coach detects your mood in real time reveals the fascinating technology behind more empathetic and responsive interactions.

Ethical AI Coaching: Ensuring Fair and Unbiased Guidance

Uncover essential strategies for ethical AI coaching that ensure fairness and unbiased guidance, and learn how to maintain integrity in your practices.

Why AI Coaching Helps Ambitious People Slow Down Intelligently

Navigate the benefits of AI coaching that helps ambitious people slow down intelligently, unlocking your potential—discover how it can transform your success journey.