
Imagine preparing a complex recipe with the most meticulous chef, yet still missing the secret ingredient that makes all the difference. In the world of AI, diligent effort isn’t enough—prioritization and strategic focus matter more than just volume of work. Just like choosing the right utensils can turn a good meal into a masterpiece, selecting the best AI models can be the key to winning or losing critical deals.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Introducing the Firmulate Experiment: Testing AI’s Business Skills in a Controlled Environment
Recently, the public platform Firmulate hosted a groundbreaking live experiment to evaluate how different AI models perform in operational decision-making—specifically, whether they can navigate a simulated company’s worst week. Four leading models, including the top-scoring GPT-5.6-sol, were tasked with managing the crisis-ridden scenario, facing identical challenges: customer issues, internal crises, and manipulation attempts.
All four models proved their competence by identifying every crisis and refusing every attempt at manipulation, such as social engineering tricks. Interestingly, only two out of the four managed to close a critical €55,000 deal. The others, despite thorough analysis, left the deal on the table, illustrating that diligent work alone doesn’t guarantee success.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deeper into Company Files
The decisive factor separating the successful models from the rest was their ability to uncover information buried two documents deep in the company’s internal files. While all models responded well to surface-level crises, those that read deeper found the key evidence necessary to close the deal at full price—adding over €4,583 in monthly recurring revenue (MRR) to the company’s bottom line. This highlights a critical insight: thoroughness isn’t just about volume; it’s about strategic depth.
As an affiliate, we earn on qualifying purchases.
Trust, Discipline, and the Risk of Slip-Ups
The most detailed participant, Opus 4.8, learned over 80 rules and conducted the deepest analyses. Yet, it still finished last because it failed to escalate certain issues and left opportunities unpursued, including a potential signature on the deal. Its discipline slipped under pressure, showing that even the most diligent AI can falter if not properly prioritized. Interestingly, this pattern of slip-ups appeared consistently across all models, albeit weaker in some.
As an affiliate, we earn on qualifying purchases.
Real-World Implications: The AI’s Role in Business and Decision-Making
For executives and managers, this experiment underscores a vital lesson: in AI-driven processes, volume of effort isn’t enough. Success depends on prioritization—reading deeply, understanding what matters most, and resisting shortcuts—especially when stakes are high. The experiment also revealed that models trained and configured without an effort parameter, such as Kimi K3, performed with the clearest discipline, closing deals reliably.
AI cybersecurity and social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering Tests: AI’s Resistance to Manipulation
In scenarios designed to test social engineering vulnerabilities, all models refused to cooperate with fake CEO messages and reporter tricks. Kimi K3 provided a clear reasoning: treating suspicious requests as possible impersonation. This resilience reinforces the importance of built-in safeguards and awareness of manipulation tactics in AI models designed for business operations.
What the Experiment Tells Us About AI and Business Readiness
The live company setup created by Firmulate is more than a simulation; it’s a mirror for real-world deployments. With over 680 self-learned rules, daily updates, and actual money mechanics, the platform demonstrates that AI models can handle crises and make sound decisions—if properly trained and focused. Yet, even the most thorough model can falter in execution, emphasizing the importance of strategic prioritization over sheer diligence.
As AI continues to integrate into customer support, CRM, forecasting, and more, decision-makers must ask: is their AI prepared to finish what it starts? Will it read the necessary information deeply enough to make informed choices? And crucially, can it resist manipulation and pressure?

The Firmulate experiment shows that diligence alone doesn’t win deals—prioritization and focus matter. AI must read deeply, stay honest, and be trained to recognize what’s truly important to succeed in real business scenarios. For companies considering AI, the lesson is clear: test your models in simulated crises before deploying them in the wild.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.