AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Good advice is easy. Acting on it is the test.

In business, the important moment often comes after the analysis: someone has to make the call, follow the rules under pressure and take responsibility for the result. That’s a question for any company considering AI agents—and one Firmulate has turned into a live experiment. Its public brand invites visitors to watch AI models run a small software company through a week of crises. The story is less about clever answers than about what happens when it’s time to act.

One company, the same hard week

Firmulate’s final Crucible League, in July 2026, put frontier models through the same customers, crises and temptations. Every decision was versioned and auditable. The published standings placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The principle behind that limit is plain: “no amount of good work outweighs a breach of trust.”

The models recognized every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap with a line that could describe plenty of workplace decisions: “Same diagnosis, same pitch — no signature.”

The clue was already in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder that business judgment can depend on finding relevant information in the material a company already has—and then carrying that insight through to a decision.

Trust faced a separate test. Fake messages claiming to come from the CEO escalated through three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the close

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The deal was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. Readers can try it at firmulate.com.

A live company, then a company-specific test

The experiment’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Visitors can watch the live experiment at firmulate.com.

That public demonstration leads to a more practical question for enterprise readers: what would an AI do with the pressures and playbooks inside your business? Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that company and produces a board report with a model ranking and weak points in its own playbooks. Nothing writes back to real systems. The point is to see how an AI workforce handles your company’s situations before asking it to handle them for real.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to wargaming

The league suggests that recognizing a crisis is only part of the job. Finding the buried clue, closing a sound deal and respecting boundaries under pressure all matter too. Firmulate makes that gap visible in a watchable public experiment, then offers enterprises a way to examine their own playbooks using a read-only export.

To discuss a pilot for your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Create Your Dream Wedding With AI Planning Software

New AI planning tool helps engaged couples manage wedding decisions, vendor quotes, and budgets without a professional planner.

Watch a Live Company Fight for Survival—With AI That Refuses to Cheat

A live, transparent experiment shows how AI models manage a small company under crisis, refusing manipulation and closing deals—yet facing ongoing financial struggles. Watch it unfold.

Revolutionize Your Sound: 9 AI Usb Microphones Leading The Market In 2026

Explore the leading AI-powered USB microphones of 2026, revolutionizing sound quality, usability, and versatility for creators and professionals alike.

Can AI Models Lead a Business to Success or Slip Up? Test Your Guessing Skills

Test your ability to spot which AI model manages a business ethically and effectively in a real-world simulation—discerning personalities that can make or break your company.