firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Test the new hire before the dinner rush

A smart oven can follow a recipe. Running a kitchen takes more: spotting trouble, handling pressure and knowing when to act. The same question faces businesses weighing AI agents for customer support, sales or operations: can a model carry out the work when the week goes wrong? Firmulate’s live experiment puts that question to the test, and its enterprise pilot offers a way to try the same kind of exercise against a company’s own business.

One company, one difficult week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced identical customers, crises and temptations. Every decision was versioned and auditable, and the experiment measured how the models managed the company rather than how polished their chat sounded.

In the final ranking, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Good diagnosis did not guarantee a finish

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap, captured in the experiment’s phrase “Same diagnosis, same pitch — no signature,” is the kind of operational weakness a chat demonstration can miss.

The decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result suggests that finding and using relevant information mattered as much as recognizing the situation.

Pressure came with a trust test

The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. The comparison comes with a fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to a company-specific pilot

The live Firmulate company makes the experiment watchable. Its 13 synthetic employees operate with real money mechanics: burn is €105k per month against €2.3k MRR, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. Readers can follow the live experiment at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made them.

For businesses, the next step is a pilot using a read-only export of their own company data. The exercise can put crisis scenarios against the company’s customers, pipeline and rules, then produce a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. That creates a chance to examine how an AI workforce might handle the pressure before entrusting it with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The league shows that spotting a crisis and refusing a bad request are only part of the job; following through on an earned opportunity matters too. An enterprise pilot can test those decisions against a read-only export of your business and show where models and playbooks falter, without writing back to live systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Season a Cast Iron Skillet Properly

Learn the simple, effective way to season your cast iron skillet. Achieve a non-stick surface and protect your cookware for generations with these practical tips.

Sur La Table Anniversary Sale Offers Up To 70% Off Including Le Creuset

Sur La Table’s anniversary sale features discounts up to 70%, including popular brands like Le Creuset and Staub. The sale runs for a limited time, attracting significant shopper interest.

The Lifetime Cookware Mindset: Buy Once, Use Forever

Discover how investing in durable cookware like cast iron and heirlooms can save money, reduce waste, and make cooking simpler for a lifetime of off-grid or manual kitchens.

How to Build Up a Natural Nonstick Layer

Learn practical steps to create a durable, natural nonstick coating on your cast iron. Discover tips, tricks, and myths about seasoning like a pro.