firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Test the new hire before the dinner rush

A smart oven can follow a recipe. Running a kitchen takes more: spotting trouble, handling pressure and knowing when to act. The same question faces businesses weighing AI agents for customer support, sales or operations: can a model carry out the work when the week goes wrong? Firmulate’s live experiment puts that question to the test, and its enterprise pilot offers a way to try the same kind of exercise against a company’s own business.

One company, one difficult week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced identical customers, crises and temptations. Every decision was versioned and auditable, and the experiment measured how the models managed the company rather than how polished their chat sounded.

In the final ranking, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Good diagnosis did not guarantee a finish

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap, captured in the experiment’s phrase “Same diagnosis, same pitch — no signature,” is the kind of operational weakness a chat demonstration can miss.

The decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result suggests that finding and using relevant information mattered as much as recognizing the situation.

Pressure came with a trust test

The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. The comparison comes with a fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to a company-specific pilot

The live Firmulate company makes the experiment watchable. Its 13 synthetic employees operate with real money mechanics: burn is €105k per month against €2.3k MRR, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. Readers can follow the live experiment at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made them.

For businesses, the next step is a pilot using a read-only export of their own company data. The exercise can put crisis scenarios against the company’s customers, pipeline and rules, then produce a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. That creates a chance to examine how an AI workforce might handle the pressure before entrusting it with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The league shows that spotting a crisis and refusing a bad request are only part of the job; following through on an earned opportunity matters too. An enterprise pilot can test those decisions against a read-only export of your business and show where models and playbooks falter, without writing back to live systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Dutch Oven Sizes: Matching Pot to Recipe

Learn how to choose the right Dutch oven size for your cooking. Discover practical tips, recent innovations, and real-world examples to match your needs.

How to Season a Cast Iron Skillet Properly

Learn the step-by-step process to season your cast iron skillet for a natural non-stick surface, rust protection, and better flavor. Practical tips included.

Can AI Managers Keep Their Promises? The Live Experiment Reveals All

A live AI management experiment tests whether frontier models can handle crises, avoid manipulation, and close deals—revealing their true management personalities and readiness.

How to Cook with Carbon Steel: A Starter Guide

Discover practical tips to start cooking with carbon steel. Learn seasoning, maintenance, and how it compares to cast iron for durable, responsive cookware.