
Imagine testing a new employee not by how many words they write but by whether they finish their tasks honestly and reliably—regardless of whether they’m told to cheat or cut corners. In the world of AI, a recent benchmark reveals a surprising truth: even the most passive, do-nothing AI scores 26 out of 100. What does this tell us about trust, diligence, and the future of AI in business?
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark That Looks Beyond Chat
Most people think of AI evaluation in terms of language skills or how convincingly it can chat. But the recent experiment conducted by Firmulate shifts the focus from chat quality to management discipline—how well these models handle complex, real-world decisions under pressure.
The Setup: Simulating a Small Software Company’s Worst Week
The experiment involved four of the leading AI frontier models, each running the same scenario: managing a small software business facing a barrage of crises, customer demands, and ethical temptations. Every decision was carefully versioned and auditable, ensuring transparency. The goal was straightforward: see if these models could identify issues, avoid manipulation, and uphold honest practices.
The Surprising Result: Even the Do-Nothing Baseline Scores 26
One key finding illustrates the benchmark’s honesty: a baseline model, which essentially does nothing, scores 26 points out of 100. This might seem puzzling—why isn’t it zero? The answer lies in how partial progress is valued. Even doing nothing counts for some points, because it at least avoids mistakes or misconduct. But the real revelation is that the score is capped; a single breach of trust, like signing a deceptive deal, completely caps the score at 26. This provides a clear, transparent floor for honest management performance.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
In practical terms, this experiment reveals that AI’s value isn’t just about how well it generates text or answers questions but about whether it can be trusted to act responsibly. For businesses increasingly relying on AI—whether in customer support, sales, or decision-making—this benchmark shows that integrity and diligence are measurable and critical.
The Hidden Weaknesses in AI Decision-Making
All four models successfully identified crises and refused manipulative tactics like fake CEO messages or reporter tricks. They all refused to sign deceptive deals. However, the models’ weaknesses emerged not in their judgment of external threats but in how they handled internal documents. Those that read deeper into the company files were able to close high-value deals, showing that reading and understanding the company’s own data is essential for trustworthiness and effectiveness.
In-Depth Discipline Versus Surface Performance
The Opus 4.8 model, which ran with the most comprehensive ruleset, consistently showed the deepest analysis, but still left deals on the table and slipped on process discipline. This indicates that thoroughness alone isn’t enough—discipline and adherence to protocols matter just as much in real management decisions.
What This Benchmark Tells Us About AI Trustworthiness
The key takeaway is that AI management performance can be measured in terms of integrity, diligence, and reliability. The fact that a do-nothing baseline scores 26 underscores that even minimal compliance and honesty are valuable, but breaches of trust—like signing a risky deal—are catastrophic and fully caps the performance score.
Transparency and Auditing Are Key
The experiment’s versioned, auditable decisions ensure that outcomes are clear and reproducible. This transparency is crucial for trust—businesses need to know if their AI systems can be held accountable, especially when ai models are making or supporting critical decisions.
Real-World Relevance: Measuring Management, Not Just Chat
For businesses, the question isn’t just whether AI writes well, but whether it can finish what it starts, read relevant documents, and stay honest under pressure. These benchmarks are a step toward evaluating AI systems in the context of actual management challenges, not just language generation.
Looking Forward: Wargaming Your AI Workforce
Firmulate’s approach enables enterprise teams to test their AI models against simulated business crises—without risking real money or reputation. By running these management wargames, companies can identify weak spots and build more disciplined, trustworthy AI systems before deployment.
See the Live Experiment in Action
Interested organizations can watch the live experiment unfold at firmulate.com/live. The platform simulates a real company, complete with 13 synthetic employees, daily decision-making, and real-money mechanics—giving managers a front-row seat to how AI models perform under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
