AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Managers Still Make It To 26 Points In Evaluation Tests on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Recent AI management tests show even the worst-performing models score at least 26 points, emphasizing the importance of trust and partial work in AI evaluation, as detailed in the original analysis. Top models approach 95, but trust breaches cap scores, raising questions about AI reliability.

AI management benchmarks conducted by Firmulate reveal that even the lowest-scoring AI managers achieve a score of 26 points, not zero. This finding challenges assumptions about AI failure, showing that partial progress is recognized and that trust breaches are the key factor limiting scores. The results matter because they shed light on AI reliability and management under stress, crucial for enterprise deployment. Insights from the original analysis can be found here.

The benchmark involved four frontier AI models managing a small software company during a simulated worst week, with the same crises, requests, and manipulations. For more context on AI management benchmarks, see this analysis. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, intentionally designed as a minimal effort, scored 26, illustrating that even minimal management efforts are recognized in scoring.

According to the benchmark’s rules, a single breach of trust caps the total score, meaning no model can achieve a perfect 100. This design aims to prevent grade inflation and emphasize integrity alongside competence. Interestingly, models that refused manipulative social engineering attempts generally scored higher, but thoroughness and follow-through did not always correlate with higher scores, as seen with Opus 4.8, which despite deep analysis, failed to follow through on closing deals or escalation protocols.

The tests also included trust challenges, such as fake CEO messages and background offers, which models mostly refused, demonstrating better handling of trust-related risks than task completion. However, models that read their own documentation and executed full processes scored notably higher, underscoring the importance of thoroughness and internal knowledge access.

At a glance
reportWhen: results announced July 2026
The developmentA new benchmark league tested AI managers under stress, revealing that even the weakest models score above zero, with trust breaches preventing perfect scores.
Why the Worst AI Managers Still Make It to 26 Points in Evaluation Tests
AI Management Benchmarks · July 2026

Why the Worst AI Managers Still Make It to 26 Points in Evaluation Tests

Recent AI management stress tests run by Firmulate reveal that even the lowest-scoring AI managers earn 26 points — never zero. Partial progress like triaging and reading emails counts, but a single breach of trust caps the total score, keeping every model below a perfect 100.

The benchmark put four frontier AI models in charge of a small software company during a simulated “worst week” of identical crises, requests, and manipulations.

26
Do-nothing baseline score
95
Top score — gpt-5.6-sol
0
Models achieving a perfect 100
4
Frontier models tested
73
Lowest model score — Opus 4.8
1
Trust breach caps total score
2026
Firmulate league launch

The Scoreboard

Simulated worst week · identical crises & manipulations
gpt-5.6-sol
95
Opus 4.8
73
Baseline (do-nothing)
26
MAXIMUM ACHIEVABLE: < 100 — a single breach of trust caps the total score by design, preventing grade inflation.

What the Tests Revealed

Findings from the July 2026 results
Finding 01 · Baseline

Partial Work Counts

Even minimal management effort — triaging tickets, reading emails — is recognized. The 26-point floor means zero scores are structurally impossible.

Finding 02 · Trust

Breaches Are Non-Negotiable

One violation of trust, such as accepting a manipulative request, caps the total score regardless of competence elsewhere. Integrity outranks output.

Finding 03 · Thoroughness

Docs-Readers Win

Models that read their own documentation and executed full processes scored notably higher than those that improvised.

Trust vs. Task Completion

Models handled manipulation better than follow-through
ChallengeModel ResponseOutcomeScore Impact
Fake CEO messagesMostly refused manipulative instructions✓ Handled wellPreserved trust score
Background offers / social engineeringGenerally refused✓ Handled wellPrevented score cap
Closing dealsOpus 4.8 analyzed deeply but failed to follow through~ MixedLost points despite insight
Escalation protocolsInconsistent execution across models~ MixedThoroughness gap
Internal documentation readingOnly some models accessed internal knowledge✗ UnderusedTop scorers relied on it

How Scoring Works

From minimal effort to the trust ceiling
1

Baseline Floor

Minimal management effort earns 26 points — zero is impossible by design.

2

Task Execution

Handling crises, reading documentation, and following processes adds points.

3

Trust Check

A single breach — accepting manipulation — caps the final total.

4

Final Score

Best case ~95. Perfect 100 is unreachable, preventing grade inflation.

Implications for Enterprises

What deployment teams should take away
Deployment

Trust Before Competence

Traditional benchmarks measure task accuracy while ignoring follow-through and integrity. For AI in management roles, trustworthiness under pressure is the gating factor — partial work counts, but breaches are hard limits.

Open Questions

Real-World Validity Unclear

It remains unclear how these scores translate to long-term deployment: adaptability, sustained trust, and unforeseen crises in complex environments still need validation. The impact of different training regimes and safeguards is also not yet understood.

Key Questions

Frequently asked · answered

Q1Why do even the worst AI managers score above zero?

The benchmark assigns a baseline of 26 points for minimal management efforts — acknowledging partial work like triaging or reading emails. Zero scores are structurally impossible.

Q2What limits AI scores in these management tests?

Trust breaches. Violating trust — for example by accepting manipulative requests — caps the maximum score regardless of other performance.

Q3How do these benchmarks affect business deployment?

They show trustworthiness, follow-through, and internal knowledge access are critical. Companies should prioritize these qualities when selecting or training AI systems.

Q4Are partial efforts like triaging considered valuable?

Yes — the scoring system rewards partial work, recognizing that even minimal efforts contribute to management and prevent complete failure.

Q5Will AI models improve their scores over time?

Future testing aims to reduce trust breaches and enhance thoroughness, potentially raising scores and reliability in enterprise settings.

Implications of Partial Progress and Trust Limits in AI Evaluation

The results highlight that AI models are evaluated not just on their ability to perform tasks but also on their trustworthiness and integrity. The cap at 26 points for minimal effort establishes a baseline for what constitutes basic management, while the maximum near 95 points shows that high performance is achievable but constrained by trust breaches. This has critical implications for enterprises relying on AI for management roles, emphasizing that partial work counts, but breaches of trust are non-negotiable barriers to top scores.

For organizations deploying AI in operational settings, these benchmarks suggest that ensuring AI systems can handle crises, read relevant documentation, and maintain integrity under pressure is vital. The scoring system encourages focus on trust and follow-through, which are often overlooked in traditional performance metrics, and underscores the need for transparency and accountability in AI decision-making processes.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks and Stress Testing

Traditional AI benchmarks primarily measure language proficiency or task accuracy, often ignoring management qualities like trustworthiness and follow-through. The Firmulate league, launched in 2026, aims to evaluate AI managers in realistic, high-stress scenarios, including crises, manipulative tactics, and trust challenges. The benchmark’s design intentionally includes a low baseline score of 26 to reflect minimal management efforts, with a maximum score near 95 for models that excel in trust and task completion.

Previous evaluations have shown that AI models can perform well in isolated tasks but struggle under pressure or when trust is tested. The 2026 league builds on these insights, emphasizing that partial progress is valuable but trust breaches are critical failures. The results from July 2026 demonstrate that even the worst models avoid outright failure, but their scores are capped by breaches of trust, highlighting the importance of integrity in AI management systems.

“The results show that even the weakest AI managers are capable of avoiding outright failure, but trust violations prevent reaching perfect scores. This underscores the central role of integrity in enterprise AI deployment.”

— Thorsten Meyer

Unclear Aspects of AI Trust and Performance Limits

It is still unclear how these scores translate to real-world AI deployment, especially regarding long-term trust, adaptability, and handling unforeseen crises. The benchmark measures specific scenarios, but broader implications for ongoing AI management in complex environments remain to be validated. Additionally, the impact of different training regimes or internal safeguards on trust and performance is not yet fully understood.

Future Developments in AI Management Benchmarking

Following the July 2026 results, further testing is expected to explore how AI models improve in trustworthiness and follow-through over time. Developers may focus on reducing trust breaches and enhancing documentation reading capabilities. Enterprises interested in deploying AI managers will likely monitor these benchmarks for evolving standards, potentially participating in live tests or pilot programs to evaluate AI readiness for critical management roles.

Key Questions

Why do even the worst AI managers score above zero?

The benchmark design assigns a baseline of 26 points for minimal management efforts, acknowledging partial work like triaging or reading emails. This prevents zero scores and reflects that even minimal effort has value.

What limits AI scores in these management tests?

The primary limit is trust breaches. If an AI model violates trust—such as accepting manipulative requests—it caps the maximum score, regardless of other performance aspects.

How do these benchmarks affect AI deployment in businesses?

They highlight that trustworthiness, follow-through, and internal knowledge access are critical for AI success in management roles. Companies should prioritize these qualities when selecting or training AI systems.

Are partial efforts like triaging considered valuable?

Yes, the scoring system rewards partial work, recognizing that even minimal efforts contribute to management and can prevent complete failure.

Will AI models improve their scores over time?

Future testing and development aim to reduce trust breaches and enhance thoroughness, potentially raising scores and reliability in enterprise settings.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Next In Colorado: The Intersection Of Supply-Chain Operations And Political Shifts

Recent developments in Colorado highlight how supply-chain operations are increasingly impacted by political shifts, with implications for trade and logistics strategies.

The End Of An Era? Europe’s AI Sector Looks Beyond Palantir

European governments are increasingly replacing Palantir with domestic and alternative systems for military and intelligence data analysis, signaling a shift in sovereignty efforts.

Technology Operations Signal Monitor: How Google Helped Destroy Adoption Of RSS Feeds (2023)

Analysis shows Google significantly contributed to the decline of RSS feed usage through platform and tooling changes in 2023.

Langenscheidt Jugendwort Abstimmung

Die Abstimmung für das Jugendwort 2026 bei Langenscheidt ist gestartet. Jugendliche können online ihre Favoriten wählen, die nächste Woche bekannt gegeben werden.