
Imagine a scenario where artificial intelligence doesn’t just handle customer queries or automate tasks, but actually manages a real company through its toughest week—making decisions, resisting temptations to cheat, and closing deals worth thousands of euros. That’s exactly what the latest live experiment by Firmulate demonstrates, revealing how AI models are stepping into the realm of genuine management and strategic decision-making.
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Testing AI in the Trenches: The Simulated Business Battle
In a groundbreaking live experiment, four leading AI models faced the same challenge: run a small software company’s worst week, complete with customer crises, internal temptations, and urgent decisions. These scenarios are designed to mimic real-world business pressures, such as handling customer service hiccups, resisting manipulation attempts, and closing lucrative deals. The goal? Measure whether these AI systems can perform as competent managers, not just chatbots or support agents.
The League of AI Competitors
The four models included gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. They competed in a fair league, with scores based on their ability to identify critical data, maintain discipline, and achieve key business outcomes. The league’s final standings were:
- gpt-5.6-sol — 95
- Kimi K3 (Moonshot) — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
Notably, all models correctly identified every crisis, refused manipulation attempts—such as fake CEO messages and reporter scams—and demonstrated decision-making integrity. Only two models successfully closed the €55,000 deal their own analysis had earned, proving that AI can, under pressure, perform complex management tasks and finalize profitable deals.
The Hidden Edge: Reading the Company’s Files
One key insight from the experiment was that the decisive advantage often lay in the models’ ability to interpret internal company documentation. The models that delved into the company’s own files—rather than just reacting to external prompts—found crucial embedded information. The winning model, Kimi K3, was able to uncover a buried reference deep within the files that clinched the deal at full price, worth an additional €4,583 in monthly recurring revenue.
Handling Social Engineering and Ethical Challenges
As part of the test, the models faced social engineering tactics—fake messages from a CEO escalating over three stages, plus a reporter trick asking for a quick on-background yes/no. All five models refused these manipulative tactics, reasoning that such requests could be impersonation or approval bypasses. Kimi K3’s rationale was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: Live, Money-Driven Operations
The experiment isn’t just theoretical. The live company involved 13 synthetic employees executing real money mechanics—burning €105,000 each month against only €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned rules. Every decision was versioned and transparent, providing a watchable window into the AI’s management style and performance at firmulate.com/live.
Lessons from the Field: Discipline and Focus Matter
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ultimately finished last. It left the close on the table and slipped discipline—writing attempts into a locked department instead of escalating—demonstrating that thoroughness alone doesn’t guarantee success. Discipline and strategic focus are crucial, especially under pressure.
Implications for Business Leaders
This experiment reveals a crucial truth: the quality of AI management agents is not solely about chat or language fluency. It’s about whether they can finish what they start, read and interpret internal documents, remain honest under manipulation, and close deals effectively. As AI continues to integrate into customer relationship management, support, and forecasting, the question shifts from “Can it write well?” to “Will it deliver results when stakes are high?”
The League Table and Fairness Note
The scores reflect the AI models’ overall management performance, with gpt-5.6-sol leading, followed closely by Kimi K3. The experiment maintained fairness by running K3 without an effort parameter (the API default), while others ran at xhigh, ensuring a level playing field.
Why This Matters for Your Business
The live leaderboard and findings underscore an essential takeaway: choosing an AI model isn’t about the highest chat scores or superficial demos. It’s about selecting one that can withstand real-world pressures, interpret critical internal data, and stay honest under stress—traits that can determine whether your AI improves operations or simply adds noise.

AI models are now capable of managing real businesses under pressure, outperforming many expectations. The key is focusing on their ability to read internal data, resist manipulation, and finish what they start—traits that matter far beyond chat quality. Firms choosing their AI workforce should prioritize performance in real management scenarios, not just demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
