
Imagine a world where your AI assistant isn’t just a chatbot but the actual boss making crucial decisions in a busy software company. Sounds like science fiction? Not anymore. Thanks to a ground-breaking live experiment, real AI models are being put through their paces in a high-stakes management simulation — and the results are revealing more than just who’s the smartest.
The Live AI Management Wargame: A New Kind of Reality Check
At the cutting edge of AI innovation, a fascinating experiment is unfolding at firmulate.com. Here, four of the world’s top frontier AI models are running a real, functioning software company — complete with customers, crises, and even money — all in a controlled environment that’s as close to real-life as it gets. This isn’t a demo or a chatbot test; it’s a full-blown management simulation designed to measure how these models handle the chaos and complexity of running a business.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Crisis, Different Minds
Every AI model faced the same challenging week: a series of customer complaints, ethical dilemmas, and opportunities to manipulate the system. Each model made decisions independently, with all actions recorded and auditable. The goal? To see if AI can not only recognize problems but also act ethically and effectively under pressure.
business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Results
- All four models identified every crisis: each recognized trouble when it appeared.
- All refused manipulation attempts: no model was tricked into cheating or cutting corners.
- Only two signed the €55,000 deal: even with the same diagnosis and pitch, only two models followed through to closing, earning full payment.
Interestingly, the decisive factor wasn’t just in the surface decisions but buried two documents deep in the company’s files — a detail that only the models that carefully read their internal data uncovered. Those who found the hidden info secured the deal at full price, adding an extra €4,583 MRR (monthly recurring revenue).
AI ethical dilemma training programs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Personalities of AI Models
What does this tell us about AI personalities? The models exhibit management styles that could be likened to human traits:
- The Detail-Oriented: The model that thoroughly read internal documents and secured the deal showed meticulousness similar to a diligent manager.
- The Terse and Focused: Another model, running at default API settings without extra effort, signed the deal quickly but missed key details.
- The Disciplined Ethical Leader: The model that refused all manipulation and suspicious requests exemplifies integrity under pressure.
Meanwhile, the most thorough participant, Opus 4.8, with over 80 learned rules, performed well but slipped on closing due to discipline lapses — leaving deals on the table instead of escalating issues properly.
As an affiliate, we earn on qualifying purchases.
Real Money, Real Risks
This experiment is unfolding in a live, functioning company with 13 synthetic employees and actual financial mechanics: burning €105k a month against a revenue of just €2.3k, with a public cash countdown and a constantly evolving set of rules. Every decision and action is part of a transparent, ongoing test accessible for watching at firmulate.com/live.
The Implications for Business and AI Trustworthiness
Beyond the thrill of seeing AI models manage a company, the core question emerges: can these systems be trusted in real-world applications? The experiment shows that, at least in this scenario, all models identified crises and rejected manipulative tactics. Yet, only some follow through on commitments and read internal data thoroughly enough to close deals.
For enterprises contemplating deploying AI in roles that demand honesty, diligence, and strategic thinking, these findings are critical. The focus shifts from just “writing well” to “finishing what’s started,” reading important data, and maintaining integrity when under pressure — qualities that are measurable in these live tests.
The Verdict: Who Comes Out on Top?
- gpt-5.6-sol: 95 points, uncovered the buried fact, closed the big deal — a complete winner.
- Kimi K3: 93 points, the new contender, signed at full price with the cleanest discipline.
- Sonnet 5: 88 points, closed the deal but with some process slips.
- Fable 5: 77 points, also closed but showed even more slips.
This leaderboard isn’t just about scores; it’s about which AI can truly handle the messy, unpredictable world of business — with honesty, thoroughness, and resolve.

This live experiment reveals that AI models can recognize crises, refuse manipulation, and even close deals — but only some read all the data and stay disciplined enough to finish the job. As AI moves closer to real management roles, these insights show what qualities matter most: honesty, diligence, and the ability to follow through under pressure. Watch the full results unfold at firmulate.com/live and see how different AI personalities shape business outcomes.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html