firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A garage can have the right diagnosis and still lose the job. A customer hesitates, a supplier changes terms, or someone sends an urgent message claiming to be the boss. When AI starts handling quotes, bookings or customer support, the question is whether it can make the right call under pressure—and follow through.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Firmulate puts that question into a live, watchable experiment. Its public brand tracks AI models running a company through crises, with real money mechanics and decisions readers can inspect.

Same company, same crises

In the final Crucible League, in July 2026, frontier models faced the same small software company and its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline result was not about spotting trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came afterward: only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. In a garage, the equivalent might be identifying what a vehicle needs, explaining the repair, then failing to secure the customer’s approval.

The answer was buried in the files

The decisive competitor weakness was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that useful business context may sit outside the conversation in front of an AI agent: in past quotes, service records, customer notes or operating rules.

The experiment also tested pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the close

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. A capable system can notice risk and still stumble over authority, persistence or the final action.

One fairness detail belongs with the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. The experiment is real and watchable at firmulate.com.

For an enterprise, the next step is a pilot built from a read-only export of its own business. The team can run crisis scenarios against that company and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. A garage group could use the exercise to see how agents handle customer pressure, operational disruptions and approval boundaries before giving them a role in day-to-day work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

A convincing demo shows what an AI can say. A company-specific wargame asks what it does when context is buried, pressure rises and the right next step requires follow-through. To discuss a Firmulate pilot using a read-only business export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management at the Crossroads: Beyond Chat Scores in Crisis Handling

AI management tests reveal crucial gaps beyond chat scores—can your AI finish what it starts, read critical files, and stay honest under pressure? Watch real crises unfold.

One-Person Brake Bleeding: The Cleanest Methods Compared

Brake bleeding alone can be clean and efficient—discover which method keeps your workspace tidy and why it matters.

Why a ‘Do-Nothing’ AI Benchmarks at 26 Points — And What It Means for Your Garage Business

Discover why even a do-nothing AI scores 26 points in benchmarks, what this reveals about trust and discipline, and how live testing ensures your automotive AI is ready for real-world challenges.

Inside a Zero-Employee Startup Battling for Survival with AI and Public Scrutiny

A live, transparent AI-managed company loses €105K/month but successfully detects crises, resists manipulation, and uncovers hidden info—watch it all unfold online.