
Imagine trusting your favorite kitchen gadget — a high-end blender or oven — to perform under pressure, to follow through on the recipe, and to stay honest when the heat is on. Now, what if that gadget was an AI managing a tiny but real software company, facing the same stress, crises, and temptations? This is exactly what Firmulate’s groundbreaking live experiment does, but with artificial intelligence instead of appliances. The results may surprise you—and could be a glimpse into how AI will manage your business, or even your kitchen someday.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Stepping into the AI’s Shoes: The Live Business Wargame
At the core of the experiment are four frontier AI models, each tasked with running a small, real software business through its worst week. From customer crises to internal temptations, every scenario is identical across all models—ensuring a fair, apples-to-apples comparison. This isn’t just a test of how well an AI can chat; it’s a measure of management integrity, problem-solving, and decision-making under pressure.
The Rules of Engagement
Every decision is versioned and auditable, meaning that each move the AI makes is stored and can be reviewed later. The models are tested for their ability to:
- Identify and resolve crises,
- Resist manipulative tactics,
- Make high-value deals,
- Follow company policies,
- Read and consider internal data,
- Maintain honesty when it’s tempting to cut corners.
Remarkably, all four models detected every crisis and refused every manipulation attempt, from fake CEO messages to reporters trying to elicit answers for a scoop.
Decisions That Matter: Signing the Big Deal
The real test was whether the models would sign a €55,000 deal based on their own analysis. Out of the four, only two models actually signed the deal—meaning they trusted their diagnosis and were able to follow through with conviction. The other two, despite making the same diagnosis and pitch, left the deal on the table, showing a lack of discipline or confidence.
The Hidden Weakness: Reading Between the Lines
Digging deeper, the experiment uncovered a crucial detail: the decisive advantage came from models that read and understood internal company documents. Those models, which examined references two layers deep into the company’s files, successfully found a critical piece of information buried in the company’s own records. This gave them a decisive edge, allowing them to close the deal at full price—an additional +€4,583 MRR (monthly recurring revenue).
Handling Social Engineering and Pressure
In a simulated social engineering attack, fake CEO messages escalated over three stages, plus a reporter’s subtle background question. All five AI models refused to be manipulated, with Kimi K3 specifically reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that even under social pressure, these models upheld integrity.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Kind of Managers Are These AIs?
Each model displays a distinct management personality. For example, Opus 4.8 proved to be the most thorough, analyzing over 80 learned rules and diving deep into issues, but ultimately left the deal unclosed, slipping into a siloed approach. Meanwhile, Kimi K3 ran without an effort parameter, focusing on fairness and discipline, and managed to sign the deal with a clean record. Sonnet 5 fell somewhere in the middle, closing the deal but with some process slips. Interestingly, all models showed their weaknesses in process discipline and focus—revealing that even advanced AI can struggle with consistency under pressure.
Why This Matters for Business and Beyond
For companies considering deploying AI in management or customer-facing roles, these findings are critical. The question is no longer just about how well an AI can generate text or handle inquiries—it’s whether it can deliver trustworthy, consistent decisions when stakes are high. The experiment’s results suggest that AI models can be trained and tested for management integrity before they are entrusted with real responsibilities.
Watch and Learn Live
The entire experiment runs live at firmulate.com/live, demonstrating a real company in real-time, every workday. Watch how decisions unfold, see the AI’s reasoning, and learn how management quality can be measured in digital form. For businesses curious to test their own AI workforce, they can try a read-only version of the same wargame, ensuring their AI’s readiness before full deployment.

AI models have distinct management personas, and in a high-stakes business scenario, only some can deliver consistent, honest decisions. The live experiment shows that AI can be tested for integrity and discipline before being trusted with real responsibilities—a game-changer for automation and management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.