AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a scenario where artificial intelligence doesn’t just handle customer queries or automate tasks, but actually manages a real company through its toughest week—making decisions, resisting temptations to cheat, and closing deals worth thousands of euros. That’s exactly what the latest live experiment by Firmulate demonstrates, revealing how AI models are stepping into the realm of genuine management and strategic decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Testing AI in the Trenches: The Simulated Business Battle

In a groundbreaking live experiment, four leading AI models faced the same challenge: run a small software company’s worst week, complete with customer crises, internal temptations, and urgent decisions. These scenarios are designed to mimic real-world business pressures, such as handling customer service hiccups, resisting manipulation attempts, and closing lucrative deals. The goal? Measure whether these AI systems can perform as competent managers, not just chatbots or support agents.

The League of AI Competitors

The four models included gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. They competed in a fair league, with scores based on their ability to identify critical data, maintain discipline, and achieve key business outcomes. The league’s final standings were:

  • gpt-5.6-sol — 95
  • Kimi K3 (Moonshot) — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

Notably, all models correctly identified every crisis, refused manipulation attempts—such as fake CEO messages and reporter scams—and demonstrated decision-making integrity. Only two models successfully closed the €55,000 deal their own analysis had earned, proving that AI can, under pressure, perform complex management tasks and finalize profitable deals.

The Hidden Edge: Reading the Company’s Files

One key insight from the experiment was that the decisive advantage often lay in the models’ ability to interpret internal company documentation. The models that delved into the company’s own files—rather than just reacting to external prompts—found crucial embedded information. The winning model, Kimi K3, was able to uncover a buried reference deep within the files that clinched the deal at full price, worth an additional €4,583 in monthly recurring revenue.

Handling Social Engineering and Ethical Challenges

As part of the test, the models faced social engineering tactics—fake messages from a CEO escalating over three stages, plus a reporter trick asking for a quick on-background yes/no. All five models refused these manipulative tactics, reasoning that such requests could be impersonation or approval bypasses. Kimi K3’s rationale was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business: Live, Money-Driven Operations

The experiment isn’t just theoretical. The live company involved 13 synthetic employees executing real money mechanics—burning €105,000 each month against only €2,300 in monthly revenue, with a public cash countdown and over 680 self-learned rules. Every decision was versioned and transparent, providing a watchable window into the AI’s management style and performance at firmulate.com/live.

Lessons from the Field: Discipline and Focus Matter

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ultimately finished last. It left the close on the table and slipped discipline—writing attempts into a locked department instead of escalating—demonstrating that thoroughness alone doesn’t guarantee success. Discipline and strategic focus are crucial, especially under pressure.

Implications for Business Leaders

This experiment reveals a crucial truth: the quality of AI management agents is not solely about chat or language fluency. It’s about whether they can finish what they start, read and interpret internal documents, remain honest under manipulation, and close deals effectively. As AI continues to integrate into customer relationship management, support, and forecasting, the question shifts from “Can it write well?” to “Will it deliver results when stakes are high?”

The League Table and Fairness Note

The scores reflect the AI models’ overall management performance, with gpt-5.6-sol leading, followed closely by Kimi K3. The experiment maintained fairness by running K3 without an effort parameter (the API default), while others ran at xhigh, ensuring a level playing field.

Why This Matters for Your Business

The live leaderboard and findings underscore an essential takeaway: choosing an AI model isn’t about the highest chat scores or superficial demos. It’s about selecting one that can withstand real-world pressures, interpret critical internal data, and stay honest under stress—traits that can determine whether your AI improves operations or simply adds noise.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

AI models are now capable of managing real businesses under pressure, outperforming many expectations. The key is focusing on their ability to read internal data, resist manipulation, and finish what they start—traits that matter far beyond chat quality. Firms choosing their AI workforce should prioritize performance in real management scenarios, not just demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Indoor Vs Outdoor Endless Pools: Which One Should You Choose?

Outdoor or indoor endless pools—discover which option best fits your lifestyle and why the decision can be more impactful than you think.

Childproofing an Endless Pool: Covers, Gates & Alarms

Just ensuring your endless pool is childproof involves key safety measures that could prevent accidents—discover how to keep your little ones safe.

Endless Pool Vs Lap Pool: Which One Fits Your Space and Goals?

Narrowing down between an endless pool and lap pool depends on your space, goals, and budget—discover which option suits your needs best.

How Loud Is an Endless Pool? The Noise Facts No One Tells You

Neither you nor your neighbors will believe how quiet an Endless Pool truly is—discover the surprising noise facts no one tells you.