firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A garage can have the right diagnosis and still lose the job. A customer hesitates, a supplier changes terms, or someone sends an urgent message claiming to be the boss. When AI starts handling quotes, bookings or customer support, the question is whether it can make the right call under pressure—and follow through.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Firmulate puts that question into a live, watchable experiment. Its public brand tracks AI models running a company through crises, with real money mechanics and decisions readers can inspect.

Same company, same crises

In the final Crucible League, in July 2026, frontier models faced the same small software company and its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline result was not about spotting trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came afterward: only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. In a garage, the equivalent might be identifying what a vehicle needs, explaining the repair, then failing to secure the customer’s approval.

The answer was buried in the files

The decisive competitor weakness was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that useful business context may sit outside the conversation in front of an AI agent: in past quotes, service records, customer notes or operating rules.

The experiment also tested pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the close

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. A capable system can notice risk and still stumble over authority, persistence or the final action.

One fairness detail belongs with the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The experiment is real and watchable at firmulate.com.

For an enterprise, the next step is a pilot built from a read-only export of its own business. The team can run crisis scenarios against that company and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. A garage group could use the exercise to see how agents handle customer pressure, operational disruptions and approval boundaries before giving them a role in day-to-day work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

A convincing demo shows what an AI can say. A company-specific wargame asks what it does when context is buried, pressure rises and the right next step requires follow-through. To discuss a Firmulate pilot using a read-only business export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

DOT 3 Vs DOT 4 Vs DOT 5.1: the Simple Guide People Wish Existed

Keen to choose the right brake fluid? Discover the simple guide to DOT 3, DOT 4, and DOT 5.1 that every driver needs.

ABS Bleeding: When You Need a Scan Tool (and When You Don’t)

Discover whether you need a scan tool for ABS bleeding and learn how to handle simple procedures or when to seek professional help.

AI’s Hidden Power: Why Getting the Deal Done Matters More Than Just Chatting About It

AI models excel at diagnosing problems, but only some finish the job—closing deals or executing tasks—when it counts. Real tests reveal true management strength.

AI Models in Action: Diligence Alone Won’t Seal the Deal — Prioritization Matters

AI models excel at spotting crises and refusing manipulation, but success depends on prioritization and disciplined action—less volume, more focus wins the race.