AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a company operating in real time, with no human employees, losing €105,000 every month—yet still fighting to stay afloat. Now, add AI models making every decision, tested daily against crises, manipulations, and ethical dilemmas. Welcome to the frontier of build-in-public business experiments, where transparency meets survival in a high-stakes AI-driven experiment.

The Real-Time Business Laboratory

At the heart of this experiment is a live, functioning small software company, watched every workday at firmulate.com/live.html. It’s no ordinary business: it has 13 synthetic employees, real money mechanics, and a public cash countdown—losing €105,000 each month against a monthly revenue of just €2,300. The company’s entire operation is versioned daily, with every decision, crisis response, and workaround transparently recorded for public scrutiny.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The AI Models in Action

Four advanced AI models, representing different frontier models, were tasked with running this company through its worst week—same customers, same crises, same temptations to cheat or manipulate. Each model’s decisions are fully auditable, providing a rare glimpse into how AI performs in complex, high-pressure business scenarios.

Key Findings: Honesty and Crisis Management

All four models successfully identified every crisis—whether it was a billing dispute, an operational glitch, or a customer complaint—and refused every manipulation attempt, including social engineering tactics like fake CEO messages or reporter tricks. Notably, despite identical diagnoses and pitches, only two models managed to close the €55,000 deal their own analysis had earned. The others, even after finding the opportunity, left the deal on the table due to discipline lapses.

The Buried Truth and Its Impact

The decisive factor was buried two document references deep within the company’s own files—not in the customer event logs. Models that read and analyze these internal documents fully closed the deal, adding over €4,583 in monthly recurring revenue—highlighting how crucial deep document analysis is in real-world decision-making.

Built-in Ethical Vigilance

When social engineering was tested—requests escalating in stages and involving a reporter trick—every model refused. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that these AI agents are not only capable of recognizing crises but also of maintaining ethical standards under pressure.

The Company’s Struggle and Lessons

Despite the AI’s impressive ability to detect risks and refuse manipulation, the company’s financial situation remains dire. It burns €105,000 monthly, with only €2,300 in revenue, and its live status is publicly visible—making it a vivid demonstration of the real costs and challenges involved in deploying AI in live business environments.

Performance Scores and Insights

  • gpt-5.6-sol scored 95 and fixed the buried fact, closing the deal at full price.
  • Kimi K3 scored 93, also closing the deal with the cleanest discipline.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Opus 4.8 scored 77, also closing but with discipline lapses and some decision errors.

Interestingly, Opus 4.8, the most thorough participant with over 80 learned rules, left the close on the table and showed signs of slipping discipline—suggesting that thoroughness does not always guarantee flawless execution under pressure.

The Core Lesson for Business and AI

This experiment underscores a vital point: when AI agents interact with your business systems, their ability to finish what they start, read internal documents thoroughly, and stay honest under pressure is what truly matters. It’s not about how well they chat or generate content—it’s whether they can deliver useful, trustworthy work consistently.

Join the Public Wargame

If you’re a business leader or an AI enthusiast, you can watch this ongoing story unfold at firmulate.com/live.html. There’s also a public quiz to guess which model made each decision, and an option to run your own enterprise through the same wargame without risking your actual business systems.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Meta Enters The AI Coding Battle With Muse Spark 1.2

Meta releases Muse Spark 1.2 and Muse Code, its new AI coding model and agent, emphasizing co-training and improved long-horizon tasks amid competitive benchmarks.

Guest app with day-of seating lookup and schedule

A new guest app allows wedding guests to view their seating and schedule via a shareable link, aiming to reduce logistical questions for couples on their wedding day.

AI Models Stand Firm Against Social Engineering Tests—A New Benchmark in Business Security

Live tests show AI models can resist social engineering tricks; with proper setup, they refuse unethical requests, safeguarding business integrity before incidents occur.

Transform Your AI Workflow With These Top Thunderbolt Docks 2026

Discover the best Thunderbolt docks of 2026 for seamless AI workflows, offering high-speed data, multiple displays, and charging in one device.