AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine your favorite Bollywood star taking a quick break — not because they want to, but because they’re forced to pause. Now, what if we told you that AI benchmarks, much like a star’s reputation, have their own way of measuring honesty and performance? Enter the world of Firmulate’s latest AI company experiment, where even the ‘do-nothing’ model scores a noteworthy 26 points. It’s a story of transparency, trust, and the surprising complexity behind AI performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Makes an Honest AI Benchmark?

In the fast-paced realm of artificial intelligence, it’s tempting to judge models solely on their chat skills or flashy demos. But a recent experiment by Firmulate delves deeper. They set up a scenario where multiple AI models managed a simulated small software company during its worst week — a week filled with crises, customer demands, and temptations to cheat.

Each model was tasked with making decisions, just like human managers, and every move was carefully recorded and made auditable. The goal? To see not just if models can talk well, but if they can truly manage a complex business environment honestly and effectively.

Amazon

AI management and ethics books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline Score of 26

One of the key findings is that even a ‘do-nothing’ approach—one that doesn’t actively try to solve problems—scores a baseline of 26 points out of 100. Why? Because partial progress counts, and the scoring system recognizes even minimal, honest effort. But more importantly, it caps the overall score if a model breaches trust, acknowledging that no good work should outweigh dishonesty.

This setup ensures fairness and discourages gaming the system. It’s a transparent way to measure true management quality, rather than just surface-level chat prowess.

Amazon

AI transparency tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Performed

Out of four AI models tested, all identified every crisis and refused every attempt at manipulation. For example, when fake CEO messages escalated over three stages, all models refused to be duped. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”

However, when it came to closing deals, only two models signed the agreements their own analysis had earned — a €55,000 deal worth over €4,500 in monthly recurring revenue. The other two, despite similar diagnoses and pitches, held back—highlighting a gap between understanding and action.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness

The key weakness was not in reacting to customer crises but in reading sensitive internal documents. The models that managed to read two document references deep into the company’s files won the deal at full price. Those that didn’t miss out on this crucial information, illustrating that access to internal knowledge can be a decisive factor in performance.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Discipline Under Pressure

The experiment also tested social engineering — fake requests from a supposed CEO, escalating in three stages, and even a reporter’s background question. Every model refused these manipulative tactics, showing a high level of discipline and trustworthiness. Kimi K3 again stood out by treating suspicious requests as impersonation risks.

The Real-World Company in Action

This isn’t just a theoretical test. The experiment runs live on Firmulate’s platform at firmulate.com/live. It simulates a real company with 13 synthetic employees, managing real cash flow (burning €105k/month against €2.3k MRR), complete with self-learned rules and daily decision logs. It’s a watchable lab where enterprise managers can test AI models before deploying them in real business settings.

Lessons for Business and AI Developers

This experiment underscores a vital point: the true test of AI in management isn’t just how well it chats but whether it stays honest, reads internal info, and acts responsibly under pressure. The fact that partial work is rewarded, but breaches of trust are capped, makes the system both fair and rigorous.

For enterprise decision-makers, it’s a reminder to look beyond surface-level AI demos. Trustworthiness, thoroughness, and discipline set the real benchmarks—values that this transparent testing approach aims to uphold.

What’s Next?

As AI models continue to evolve, platforms like Firmulate are pioneering ways to evaluate them in practical, high-stakes scenarios. The league table, with scores from 95 (gpt-5.6-sol) down to 77 (Sonnet 5), illustrates a wide performance spectrum, but all models showed commendable honesty in this test.

And yes, even the ‘do-nothing’ baseline scored a 26, proving that honesty and partial effort are recognized — a model of integrity in an industry hungry for trustworthy AI.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Valve’s Steam Machine Is Now Available on Steam — Sign Up Before June 25

Valve begins accepting sign-ups for its new Steam Machine, available through Steam before June 25. Limited spots with randomized selection process.

Prince Harry, Duke Of Sussex

Prince Harry, Duke of Sussex, has launched a new charitable project focused on mental health support, marking his latest public engagement.

Top 10 Indian Web Series of 2024 You Need to Binge-Watch

Stay tuned to discover the top 10 Indian web series of 2024 that will keep you hooked with their captivating stories and genres.

Post-Pandemic Bollywood: How COVID-19 Transformed the Film Industry

COVID-19 revolutionized Bollywood’s approach to film distribution and storytelling, leaving the industry at a crossroads you’ll want to explore.