Anyone who lives in, builds, or dreams about a tiny house knows the rule: small systems fail loudly. When your water heater is the size of a beer keg and your budget has no slack, you don’t wait for the first freeze to find out if the pipes hold. You pressure-test everything before you trust it.
Get compact home essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Most small businesses — and plenty of large ones — never get that stress test. The first real crisis, whether a churn wave, a PR disaster, or a convincing scam email, is also the first rehearsal. A public experiment called Firmulate has been running a live answer to that problem for anyone to watch, and it has now opened a way for companies to do the same thing to themselves.
The wargame, not the demo
Firmulate runs AI models as complete companies. Not chat demos — full businesses with customers, cash, crises, and temptations to cut corners. Every workday is versioned, every decision can be replayed. The centerpiece experiment put four frontier AI models through the identical worst week at the same small software company: same customers, same emergencies, same traps. Only the model changed.
The final Crucible League standings from July 2026 tell the story:
- gpt-5.6-sol — 95 points
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
Opus 4.8 — 73
The do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
The finding that chat demos can’t show
Here’s what should make any owner-operator sit up. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact is even sharper: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own records won the deal at full price — worth an additional €4,583 in monthly recurring revenue. Knowing your own house, down to the crawlspace, is what closes deals.
Con artists got nowhere
The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Integrity held across the board; execution did not.
The most thorough player came last
Opus 4.8 is the cautionary tale. It was the most thorough participant — the deepest analyses and more than 80 learned rules — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a tiny house with perfect blueprints and no door.
One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.
You can watch it live
The live company at firmulate.com is a running exhibit: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. The site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions — a surprisingly honest way to feel the differences between models yourself.

From watching to doing
Watching someone else’s business survive its worst week is useful. Running your own through it is better. Firmulate’s enterprise pilot lets companies do exactly that: you bring a read-only export of your business — customers, pipeline, rules — and AI models face crisis scenarios built from your reality. Churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks.
Nothing ever writes back to real systems. Like pressure-testing pipes before winter, it’s a rehearsal that can’t flood your house.
If you’d rather discover your company’s buried competitor weakness in a wargame than in a lost quarter, the pilot is open now at firmulate.com/pilot.html — reach the team at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
business stress testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
