
Anyone who lives in a tiny house knows the truth about gauges. Your water tank sensor doesn’t say “full” or “empty” — it tells you how much you actually have, because partial matters. 30% of a tank is not 0%, and pretending otherwise gets you through a winter or leaves you stranded. Most technology reporting never learned this lesson. AI benchmarks in particular love clean, dramatic numbers: perfect scores, zero-shot failures, leaderboards that read like horse races. So it was refreshing to stumble across an experiment that treats measurement the way an off-grid homeowner treats a battery monitor — with granular honesty, and a healthy suspicion of anything that reads a suspiciously round 100.
Get compact home essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The experiment is Firmulate, and it ran something unusual: four frontier AI models, each given the same job — running the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the management equivalent of a logbook you can actually inspect.
Why a do-nothing manager still scores 26
The most interesting number in the whole exercise isn’t at the top of the league table — it’s at the bottom. A do-nothing baseline run, where the AI essentially sits on its hands, scores 26 points, not 0. That design choice deserves attention, because it encodes a real truth about management: doing nothing is rarely the worst outcome, and partial progress is real progress. A manager who handles half the inbox, prevents half the fires and closes nothing still delivered something. The benchmark refuses to flatter itself with dramatic zeros.
But the floor has a ceiling-shaped counterpart. A single breach of trust caps the total grade entirely. The framing is blunt: no amount of good work outweighs a breach of trust. In a tiny house, that’s the logic of your load-bearing wall — everything else can be imperfect, but that one failure sinks the structure. The benchmark treats honesty the same way: not as one line item to be averaged away, but as the thing that bounds everything else.
The final league
As of the final July 2026 standings: gpt-5.6-sol took first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 at 77, and Opus 4.8 last at 73. Notably, nothing hit 100 — and the experiment seems fine with that. A perfect round score would be a red flag, not a triumph.
The buried fact
Here’s the finding that should worry anyone delegating work to AI agents. All the models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Why? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The ones that didn’t read before acting left the close on the table. It’s the AI equivalent of not checking your own rainwater catchment records before sizing a new tank.
Pressure tests and a fairness footnote
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably plain: “Treat the request as a suspected approval-bypass / possible impersonation.” One caveat the experiment itself discloses: K3 ran without an effort parameter while competitors ran at high effort settings — a transparency note you rarely see in this space.
The Opus 4.8 profile is the cautionary tale: it was the most thorough participant, accumulating over 80 learned rules and the deepest analyses, yet finished last. The close went unsigned, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort without follow-through is just expensive motion.
Still running, still watchable
Firmulate didn’t stop at the tournament. It now runs a live company with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live — a slow-burn soap opera with a balance sheet. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The lesson translates cleanly from server racks to small cabins: honest measurement means counting partial progress, refusing to hand out round perfects, and treating trust as the non-negotiable structural element. A gauge that only says “good” or “bad” isn’t a gauge. Whether you’re monitoring a 200-amp-hour battery bank or an AI agent with access to your CRM, you want the instrument that tells you exactly how much you have — including when the tank reads 26.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
