AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine testing a new employee with a polished interview — only to find they fall apart when actually doing the job. That’s the gap AI models reveal in real-world scenarios. While they can mimic understanding in chat, their true test comes in decision-making under pressure. When AI is entrusted with critical business tasks, surface-level chat prowess isn’t enough — resilience, thoroughness, and honesty matter much more.

What Happens When AI Runs a Business Crisis?

Recently, a groundbreaking experiment by Firmulate exposed the true capabilities — and limitations — of top AI models in a high-stakes business simulation. Four AI models, all competing to manage a small software company through its worst week, faced identical crises, temptations, and decision points. The goal? See which AI could not only identify problems but also follow through on their analysis and close a crucial €55,000 deal.

All four models—gpt-5.6-sol 95, Kimi K3, Sonnet 5, and Fable 5—were able to spot every crisis and refused every manipulation attempt, including sophisticated social engineering tactics. That’s the surface level: in a chat demo, they all looked capable. However, the real story emerged in their actual decision-making process.

The Hidden Weakness: Reading the Files that Matter

The decisive factor was reading deeper into the company’s own documentation. While all models diagnosed issues correctly, only two—gpt-5.6-sol and Kimi K3—looked into critical internal files that contained the key to closing the deal. This buried information was two document references deep in the company’s files, not immediately visible in a chat interface. When the models accessed and understood this data, they signed the €55,000 deal, earning full revenue (+€4,583 MRR).

Discipline and Decision Quality Under Pressure

The experiment didn’t stop at crisis detection. It tested the models’ discipline under social engineering, where fake CEO messages escalated over three stages, and even a reporter’s subtle request to bypass approval. All five models refused these manipulative tactics, demonstrating strong resistance to deception. Kimi K3, in particular, justified its refusal by treating the requests as potential impersonation attempts.

Yet, despite this resilience, only two models actually completed their work — the other two, including Opus 4.8, faltered at the final hurdle. Opus, the most thorough participant with over 80 learned rules, ultimately left the deal unexecuted, slipping into process slips and attempting to write notes into a locked department instead of escalating.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Test Is Not Chat, but Execution

This experiment underscores a vital lesson: chat demos measure surface-level language skills, not real-world business execution. A model that looks impressive in a demo might still fail to finish what it starts. The ability to read your company’s internal documents, resist manipulation, and follow through on commitments is what distinguishes a reliable AI from an unreliable one.

At Firmulate, this approach is not theoretical. The live company runs every day with 13 synthetic employees, handling real money mechanics—a burn rate of €105k per month against a modest €2.3k MRR. Every decision is versioned and auditable, and the entire process happens in real-time at firmulate.com/live.

The League Table: Who Managed the Job?

  • gpt-5.6-sol 95: Found the buried fact and closed the deal — overall top performer.
  • Kimi K3 93: Closed the deal too, with the cleanest discipline among competitors.
  • Sonnet 5 88: Also closed the deal, but with minor slips.
  • Fable 5 77: Demonstrated strong rule adherence but left the deal unexecuted.

What’s clear is that scoring well on chat demos doesn’t guarantee a model’s ability to close, stick to protocols, or read the critical internal data. These are the skills that matter when AI is entrusted with real business responsibilities.

The Implication for Business Leaders

If you plan to integrate AI into customer support, CRM, or forecasting, ask yourself: does it just talk well, or can it actually finish the job? The experiment from Firmulate proves that true management quality lies in resilience, thoroughness, and integrity — qualities that are invisible in chat previews but critical in real work.

To explore how your AI workforce performs under pressure, consider running your own wargame against a read-only export of your business. It’s a safe way to test without risking real systems or money, and to see whether your AI can truly deliver on what it promises.

Learn More and Watch the Live Company in Action

Discover full results and plain-language insights at firmulate.com/benchmarks.html. To see the experiment in real-time, visit firmulate.com/live, where the company’s daily decisions unfold with transparency. Or test your own management skills with the interactive quiz at firmulate.com/quiz.html.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Surface chat skills are no guarantee of real-world AI performance. Resilience, thoroughness, and honesty under pressure are key — and only real tests reveal true capabilities. Watch the live experiment and see which models can finish the job, not just talk about it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI Models Stand Firm Against Social Engineering Tests—A New Benchmark in Business Security

Live tests show AI models can resist social engineering tricks; with proper setup, they refuse unethical requests, safeguarding business integrity before incidents occur.

Right-sized planning checklist for 30-guest weddings

A new scaled-down wedding planning checklist for 30-guest ceremonies is being tested to simplify planning for intimate weddings, addressing gaps in current tools.

Meta Enters The AI Coding Battle With Muse Spark 1.2

Meta releases Muse Spark 1.2 and Muse Code, its new AI coding model and agent, emphasizing co-training and improved long-horizon tasks amid competitive benchmarks.

2026’S Top External GPU Choices For AI Innovation

Discover the best external GPUs for AI innovation in 2026, highlighting top models, compatibility, performance, and what to consider for future-proofing.