AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine watching a high-stakes race where a newcomer suddenly overtakes seasoned contenders—only this time, the race is a test of artificial intelligence managing a real company’s worst week. In the fast-evolving world of AI, the question isn’t just about how well these models chat, but whether they can make reliable, honest decisions under pressure. The recent challenge, hosted by Firmulate, offers a revealing glimpse into this new frontier, showing that even newcomers can outshine established giants when it comes to trustworthy management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Challenge: Simulating Chaos in a Real Business

In July 2026, four frontier AI models were put through a grueling test designed to mimic the worst week a small software company could face. The scenario was crafted to include the same customers, crises, and temptations—yet only the decision-making powers of the AI models changed. This rigorous experiment ensured an apples-to-apples comparison, with every decision versioned and auditable to ensure transparency and fairness.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: A Surprising Leader Emerges

The results are striking. The models scored from 73 to 95 in a comprehensive league table, with the top scorers being gpt-5.6-sol at 95 points and Kimi K3 at 93. The latter, a newcomer from Moonshot, not only found the buried security ‘needle’ in the company’s files but also secured a €55,000 deal, increasing monthly recurring revenue (MRR) by €4,583. Meanwhile, other models, despite diagnosing correctly and pitching well, failed to sign the deal, illustrating a crucial gap between assessment and execution.

Winning Strategies and Disappointments

The K3 model demonstrated exceptional discipline, resisting all three social engineering attempts—fake CEO messages and reporter tricks—that tried to manipulate decisions. It treated suspicious requests as potential impersonation and maintained integrity. Meanwhile, Opus 4.8, a model with the most thorough analysis—over 80 learned rules—struggled with closing the deal, leaving the opportunity on the table and slipping on discipline. This highlights that thoroughness alone doesn’t guarantee success; disciplined execution under pressure is key.

The Hidden Weakness: Read-Deep Is Key

One of the most revealing findings was that the decisive advantage came from reading deeply into the company’s internal files—two references deep—rather than just reacting to customer events. Models that could access and interpret internal data at depth gained a significant edge, successfully closing the deal at full price. This underscores that in real management, understanding the full context is often more critical than surface-level crisis recognition.

The Live Experiment: Real Money, Real Consequences

This isn’t just a simulation. The live company, with 13 synthetic employees and real monetary mechanics, runs every business day at Firmulate. The company burns €105k each month against only €2.3k MRR, with a public cash countdown. Every decision made by these models is logged and observable, providing an unprecedented window into AI-driven management—where the focus is on whether the AI can truly finish what it started, stay honest, and deliver useful work, not just generate convincing chat.

The Takeaway: Trustworthiness Trumps Talk

This experiment exposes an important truth for businesses considering AI management tools. The capacity to identify critical buried data, resist manipulative tactics, and actually close deals under pressure sets the best models apart. The top-scoring K3 model achieved a perfect balance—finding hidden security issues, refusing manipulation, and completing the deal—showing that AI’s utility depends on discipline and depth, not just superficial performance.

Why This Matters for Your Business

If AI agents will someday touch your CRM, support systems, or forecasts, the concern isn’t just their conversational skill. It’s whether they can finish complex tasks reliably, stay honest under stress, and read your internal files thoroughly. The league table from this test reveals a clear gap: even among frontier models, performance varies significantly, and the best are those that combine deep data access with disciplined decision-making.

Final Thoughts: The League Is Open

As the AI management league heats up, the message is clear: choosing a model without your own rigorous testing is a gamble. The recent experiment shows that a newcomer—Kimi K3—can outperform well-established models when tested in the crucible of real management. For enterprises, this means adopting and testing AI tools in real scenarios is more crucial than ever to ensure trustworthy, effective performance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent AI management experiment shows a newcomer from Moonshot outperformed established frontier models by reliably finding buried data, resisting manipulation, and closing deals. Trust and discipline are key—test your AI before deploying.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Guest app with day-of seating lookup and schedule

A new guest app allows wedding guests to view their seating and schedule via a shareable link, aiming to reduce logistical questions for couples on their wedding day.

When Diligence Isn’t Enough: How AI’s Volume Overlooks Critical Impact in Business Decisions

Firmulate’s live AI experiment shows that thoroughness alone isn’t enough—prioritization, discipline, and impact-focused decision-making are key to AI success in real business scenarios.

Why a Do-Nothing AI Benchmarks at 26 — and What It Tells Us About Trust in AI

A recent AI benchmark shows a do-nothing model scores 26, revealing that trust, discipline, and honest management are critical. Firms can test AI in simulated business crises before deployment.

Revolutionize Your Sound: 9 AI Usb Microphones Leading The Market In 2026

Explore the leading AI-powered USB microphones of 2026, revolutionizing sound quality, usability, and versatility for creators and professionals alike.