🔍 Read the full analysis: Why The Worst AI Managers Still Make It To 26 Points In Evaluation Tests on ThorstenMeyerAI.com
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Recent AI management tests show even the worst-performing models score at least 26 points, emphasizing the importance of trust and partial work in AI evaluation, as detailed in the original analysis. Top models approach 95, but trust breaches cap scores, raising questions about AI reliability.
AI management benchmarks conducted by Firmulate reveal that even the lowest-scoring AI managers achieve a score of 26 points, not zero. This finding challenges assumptions about AI failure, showing that partial progress is recognized and that trust breaches are the key factor limiting scores. The results matter because they shed light on AI reliability and management under stress, crucial for enterprise deployment. Insights from the original analysis can be found here.
The benchmark involved four frontier AI models managing a small software company during a simulated worst week, with the same crises, requests, and manipulations. For more context on AI management benchmarks, see this analysis. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, intentionally designed as a minimal effort, scored 26, illustrating that even minimal management efforts are recognized in scoring.
According to the benchmark’s rules, a single breach of trust caps the total score, meaning no model can achieve a perfect 100. This design aims to prevent grade inflation and emphasize integrity alongside competence. Interestingly, models that refused manipulative social engineering attempts generally scored higher, but thoroughness and follow-through did not always correlate with higher scores, as seen with Opus 4.8, which despite deep analysis, failed to follow through on closing deals or escalation protocols.
The tests also included trust challenges, such as fake CEO messages and background offers, which models mostly refused, demonstrating better handling of trust-related risks than task completion. However, models that read their own documentation and executed full processes scored notably higher, underscoring the importance of thoroughness and internal knowledge access.
Why the Worst AI Managers Still Make It to 26 Points in Evaluation Tests
Recent AI management stress tests run by Firmulate reveal that even the lowest-scoring AI managers earn 26 points — never zero. Partial progress like triaging and reading emails counts, but a single breach of trust caps the total score, keeping every model below a perfect 100.
The benchmark put four frontier AI models in charge of a small software company during a simulated “worst week” of identical crises, requests, and manipulations.
The Scoreboard
What the Tests Revealed
Partial Work Counts
Even minimal management effort — triaging tickets, reading emails — is recognized. The 26-point floor means zero scores are structurally impossible.
Breaches Are Non-Negotiable
One violation of trust, such as accepting a manipulative request, caps the total score regardless of competence elsewhere. Integrity outranks output.
Docs-Readers Win
Models that read their own documentation and executed full processes scored notably higher than those that improvised.
Trust vs. Task Completion
| Challenge | Model Response | Outcome | Score Impact |
|---|---|---|---|
| Fake CEO messages | Mostly refused manipulative instructions | ✓ Handled well | Preserved trust score |
| Background offers / social engineering | Generally refused | ✓ Handled well | Prevented score cap |
| Closing deals | Opus 4.8 analyzed deeply but failed to follow through | ~ Mixed | Lost points despite insight |
| Escalation protocols | Inconsistent execution across models | ~ Mixed | Thoroughness gap |
| Internal documentation reading | Only some models accessed internal knowledge | ✗ Underused | Top scorers relied on it |
How Scoring Works
Baseline Floor
Minimal management effort earns 26 points — zero is impossible by design.
Task Execution
Handling crises, reading documentation, and following processes adds points.
Trust Check
A single breach — accepting manipulation — caps the final total.
Final Score
Best case ~95. Perfect 100 is unreachable, preventing grade inflation.
Implications for Enterprises
Trust Before Competence
Traditional benchmarks measure task accuracy while ignoring follow-through and integrity. For AI in management roles, trustworthiness under pressure is the gating factor — partial work counts, but breaches are hard limits.
Real-World Validity Unclear
It remains unclear how these scores translate to long-term deployment: adaptability, sustained trust, and unforeseen crises in complex environments still need validation. The impact of different training regimes and safeguards is also not yet understood.
Key Questions
Q1Why do even the worst AI managers score above zero?
The benchmark assigns a baseline of 26 points for minimal management efforts — acknowledging partial work like triaging or reading emails. Zero scores are structurally impossible.
Q2What limits AI scores in these management tests?
Trust breaches. Violating trust — for example by accepting manipulative requests — caps the maximum score regardless of other performance.
Q3How do these benchmarks affect business deployment?
They show trustworthiness, follow-through, and internal knowledge access are critical. Companies should prioritize these qualities when selecting or training AI systems.
Q4Are partial efforts like triaging considered valuable?
Yes — the scoring system rewards partial work, recognizing that even minimal efforts contribute to management and prevent complete failure.
Q5Will AI models improve their scores over time?
Future testing aims to reduce trust breaches and enhance thoroughness, potentially raising scores and reliability in enterprise settings.
Implications of Partial Progress and Trust Limits in AI Evaluation
The results highlight that AI models are evaluated not just on their ability to perform tasks but also on their trustworthiness and integrity. The cap at 26 points for minimal effort establishes a baseline for what constitutes basic management, while the maximum near 95 points shows that high performance is achievable but constrained by trust breaches. This has critical implications for enterprises relying on AI for management roles, emphasizing that partial work counts, but breaches of trust are non-negotiable barriers to top scores.
For organizations deploying AI in operational settings, these benchmarks suggest that ensuring AI systems can handle crises, read relevant documentation, and maintain integrity under pressure is vital. The scoring system encourages focus on trust and follow-through, which are often overlooked in traditional performance metrics, and underscores the need for transparency and accountability in AI decision-making processes.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks and Stress Testing
Traditional AI benchmarks primarily measure language proficiency or task accuracy, often ignoring management qualities like trustworthiness and follow-through. The Firmulate league, launched in 2026, aims to evaluate AI managers in realistic, high-stress scenarios, including crises, manipulative tactics, and trust challenges. The benchmark’s design intentionally includes a low baseline score of 26 to reflect minimal management efforts, with a maximum score near 95 for models that excel in trust and task completion.
Previous evaluations have shown that AI models can perform well in isolated tasks but struggle under pressure or when trust is tested. The 2026 league builds on these insights, emphasizing that partial progress is valuable but trust breaches are critical failures. The results from July 2026 demonstrate that even the worst models avoid outright failure, but their scores are capped by breaches of trust, highlighting the importance of integrity in AI management systems.
“The results show that even the weakest AI managers are capable of avoiding outright failure, but trust violations prevent reaching perfect scores. This underscores the central role of integrity in enterprise AI deployment.”
— Thorsten Meyer
Unclear Aspects of AI Trust and Performance Limits
It is still unclear how these scores translate to real-world AI deployment, especially regarding long-term trust, adaptability, and handling unforeseen crises. The benchmark measures specific scenarios, but broader implications for ongoing AI management in complex environments remain to be validated. Additionally, the impact of different training regimes or internal safeguards on trust and performance is not yet fully understood.
Future Developments in AI Management Benchmarking
Following the July 2026 results, further testing is expected to explore how AI models improve in trustworthiness and follow-through over time. Developers may focus on reducing trust breaches and enhancing documentation reading capabilities. Enterprises interested in deploying AI managers will likely monitor these benchmarks for evolving standards, potentially participating in live tests or pilot programs to evaluate AI readiness for critical management roles.
Key Questions
Why do even the worst AI managers score above zero?
The benchmark design assigns a baseline of 26 points for minimal management efforts, acknowledging partial work like triaging or reading emails. This prevents zero scores and reflects that even minimal effort has value.
What limits AI scores in these management tests?
The primary limit is trust breaches. If an AI model violates trust—such as accepting manipulative requests—it caps the maximum score, regardless of other performance aspects.
How do these benchmarks affect AI deployment in businesses?
They highlight that trustworthiness, follow-through, and internal knowledge access are critical for AI success in management roles. Companies should prioritize these qualities when selecting or training AI systems.
Are partial efforts like triaging considered valuable?
Yes, the scoring system rewards partial work, recognizing that even minimal efforts contribute to management and can prevent complete failure.
Will AI models improve their scores over time?
Future testing and development aim to reduce trust breaches and enhance thoroughness, potentially raising scores and reliability in enterprise settings.
Source: ThorstenMeyerAI.com
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
