AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Tiny house people know the difference between a good blueprint and a good build. The plans look perfect on Instagram — every joist labeled, every inch accounted for — but the truth shows up on the day the trailer arrives two feet short, the inspector wants a change, and the budget has exactly one surprise left in it. Building small is really a test of judgment under pressure, not drafting skill.

A live experiment at Firmulate just showed that the exact same gap exists in AI — and almost nobody is measuring it.

Chat quality versus management quality

Most AI rankings today — coding leaderboards, chat arenas — measure how well a model answers a question. Firmulate measures something else: how well it runs a company through its worst week. Four frontier AI models were each handed the same small software business and the same seven days of chaos: a churn wave, a price increase, a downround scare, a PR crisis. Same customers, same temptations. Every decision versioned and auditable.

The final Crucible League standings from July 2026: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.

The finding that chat demos can’t show

Here’s the headline result: all four models spotted every crisis, and all four refused every manipulation attempt. But only two actually finished the job — signing the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.

The decisive fact was buried two document references deep in the company’s own files — not in the customer event at all. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. That’s the AI equivalent of a builder who never opened the site survey: brilliant answers, house on the wrong lot.

Then there was the social engineering test: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Effort isn’t everything

Opus 4.8 is the cautionary tale for anyone who equates thoroughness with competence. It was the most diligent participant — 80 learned rules added, the deepest analyses of the field — and it still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models.

One fairness note: K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still nearly won.

You can watch it lose money

This isn’t a slide deck. The company is real software with 13 synthetic employees and real money mechanics — €105k monthly burn against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It runs live at firmulate.com, rebuilding itself twice a day.

There’s also a quiz built from 242 real, unedited management decisions, where you guess which model made which call — a humbling exercise. Enterprises can go further and run the same wargame against a read-only export of their own business; nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files first, and stay honest when the pressure is on. Tiny house builders learned long ago that the pretty plan isn’t the house. Firmulate is one of the first places testing whether AI can tell the difference — before you hand it the keys.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Guest app with day-of seating lookup and schedule

A new mobile guest app allows couples to share real-time seating and schedule info via a single link, reducing last-minute questions on wedding day.

The Radical Small-Footprint Business Where AI Workers Race a Public Cash Clock

Firmulate’s 13 synthetic employees run a cash-strapped software company in public, revealing whether AI can turn sound judgment into finished work.

Could You Spot the AI Boss? A Live Wargame Tests Management Instincts

A live business wargame reveals distinct AI management personalities—and asks readers to guess which frontier model made each real decision.

What AI Tests Reveal About Trust and Performance in Business — Not Just Chat

Real AI testing reveals that trustworthiness and follow-through matter most. Discover how models perform under pressure in a live business experiment—beyond chat demos.