AIThis post was created with the assistance of artificial intelligence (AI).

Anyone who lives in, builds, or dreams about a tiny house knows the rule: small systems fail loudly. When your water heater is the size of a beer keg and your budget has no slack, you don’t wait for the first freeze to find out if the pipes hold. You pressure-test everything before you trust it.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get compact home essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Most small businesses — and plenty of large ones — never get that stress test. The first real crisis, whether a churn wave, a PR disaster, or a convincing scam email, is also the first rehearsal. A public experiment called Firmulate has been running a live answer to that problem for anyone to watch, and it has now opened a way for companies to do the same thing to themselves.

The wargame, not the demo

Firmulate runs AI models as complete companies. Not chat demos — full businesses with customers, cash, crises, and temptations to cut corners. Every workday is versioned, every decision can be replayed. The centerpiece experiment put four frontier AI models through the identical worst week at the same small software company: same customers, same emergencies, same traps. Only the model changed.

The final Crucible League standings from July 2026 tell the story:

  • gpt-5.6-sol — 95 points
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

The finding that chat demos can’t show

Here’s what should make any owner-operator sit up. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact is even sharper: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own records won the deal at full price — worth an additional €4,583 in monthly recurring revenue. Knowing your own house, down to the crawlspace, is what closes deals.

Con artists got nowhere

The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Integrity held across the board; execution did not.

The most thorough player came last

Opus 4.8 is the cautionary tale. It was the most thorough participant — the deepest analyses and more than 80 learned rules — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a tiny house with perfect blueprints and no door.

One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.

You can watch it live

The live company at firmulate.com is a running exhibit: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. The site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions — a surprisingly honest way to feel the differences between models yourself.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to doing

Watching someone else’s business survive its worst week is useful. Running your own through it is better. Firmulate’s enterprise pilot lets companies do exactly that: you bring a read-only export of your business — customers, pipeline, rules — and AI models face crisis scenarios built from your reality. Churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks.

Nothing ever writes back to real systems. Like pressure-testing pipes before winter, it’s a rehearsal that can’t flood your house.

If you’d rather discover your company’s buried competitor weakness in a wargame than in a lost quarter, the pilot is open now at firmulate.com/pilot.html — reach the team at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business stress testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Small Systems, Honest Gauges: What a Tiny House Owner Can Learn From an AI Benchmark That Refuses to Give Out Zeros (or Perfect 100s)

An AI benchmark where doing nothing scores 26, one breach of trust caps everything, and no model gets a round 100 — honesty in measurement, tiny-house style.

The AI Boss Might Ace the Interview and Still Burn Down the Business

Four AI models each ran the same company through its worst week. All spotted every crisis — only two closed the deal. Firmulate measures management, not chat.

Small Footprint, Big Results: What a Tiny Company Simulator Just Revealed About AI

A public wargame ran five AI models as the same struggling company. Moonshot’s newcomer Kimi K3 took second place with the cleanest discipline — beating three of four Western frontier models.

The Solo Operator’s AI Test: Reading the Fine Print Before It Signs

An AI experiment buried a €55,000 fact two documents deep. Only two frontier models found it — and the gap decided the deal. Here’s why it matters.