
Imagine your favorite Bollywood star taking a quick break — not because they want to, but because they’re forced to pause. Now, what if we told you that AI benchmarks, much like a star’s reputation, have their own way of measuring honesty and performance? Enter the world of Firmulate’s latest AI company experiment, where even the ‘do-nothing’ model scores a noteworthy 26 points. It’s a story of transparency, trust, and the surprising complexity behind AI performance.
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Makes an Honest AI Benchmark?
In the fast-paced realm of artificial intelligence, it’s tempting to judge models solely on their chat skills or flashy demos. But a recent experiment by Firmulate delves deeper. They set up a scenario where multiple AI models managed a simulated small software company during its worst week — a week filled with crises, customer demands, and temptations to cheat.
Each model was tasked with making decisions, just like human managers, and every move was carefully recorded and made auditable. The goal? To see not just if models can talk well, but if they can truly manage a complex business environment honestly and effectively.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline Score of 26
One of the key findings is that even a ‘do-nothing’ approach—one that doesn’t actively try to solve problems—scores a baseline of 26 points out of 100. Why? Because partial progress counts, and the scoring system recognizes even minimal, honest effort. But more importantly, it caps the overall score if a model breaches trust, acknowledging that no good work should outweigh dishonesty.
This setup ensures fairness and discourages gaming the system. It’s a transparent way to measure true management quality, rather than just surface-level chat prowess.
As an affiliate, we earn on qualifying purchases.
How the Models Performed
Out of four AI models tested, all identified every crisis and refused every attempt at manipulation. For example, when fake CEO messages escalated over three stages, all models refused to be duped. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
However, when it came to closing deals, only two models signed the agreements their own analysis had earned — a €55,000 deal worth over €4,500 in monthly recurring revenue. The other two, despite similar diagnoses and pitches, held back—highlighting a gap between understanding and action.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness
The key weakness was not in reacting to customer crises but in reading sensitive internal documents. The models that managed to read two document references deep into the company’s files won the deal at full price. Those that didn’t miss out on this crucial information, illustrating that access to internal knowledge can be a decisive factor in performance.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Discipline Under Pressure
The experiment also tested social engineering — fake requests from a supposed CEO, escalating in three stages, and even a reporter’s background question. Every model refused these manipulative tactics, showing a high level of discipline and trustworthiness. Kimi K3 again stood out by treating suspicious requests as impersonation risks.
The Real-World Company in Action
This isn’t just a theoretical test. The experiment runs live on Firmulate’s platform at firmulate.com/live. It simulates a real company with 13 synthetic employees, managing real cash flow (burning €105k/month against €2.3k MRR), complete with self-learned rules and daily decision logs. It’s a watchable lab where enterprise managers can test AI models before deploying them in real business settings.
Lessons for Business and AI Developers
This experiment underscores a vital point: the true test of AI in management isn’t just how well it chats but whether it stays honest, reads internal info, and acts responsibly under pressure. The fact that partial work is rewarded, but breaches of trust are capped, makes the system both fair and rigorous.
For enterprise decision-makers, it’s a reminder to look beyond surface-level AI demos. Trustworthiness, thoroughness, and discipline set the real benchmarks—values that this transparent testing approach aims to uphold.
What’s Next?
As AI models continue to evolve, platforms like Firmulate are pioneering ways to evaluate them in practical, high-stakes scenarios. The league table, with scores from 95 (gpt-5.6-sol) down to 77 (Sonnet 5), illustrates a wide performance spectrum, but all models showed commendable honesty in this test.
And yes, even the ‘do-nothing’ baseline scored a 26, proving that honesty and partial effort are recognized — a model of integrity in an industry hungry for trustworthy AI.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
