AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who lives in a tiny house knows the truth about gauges. Your water tank sensor doesn’t say “full” or “empty” — it tells you how much you actually have, because partial matters. 30% of a tank is not 0%, and pretending otherwise gets you through a winter or leaves you stranded. Most technology reporting never learned this lesson. AI benchmarks in particular love clean, dramatic numbers: perfect scores, zero-shot failures, leaderboards that read like horse races. So it was refreshing to stumble across an experiment that treats measurement the way an off-grid homeowner treats a battery monitor — with granular honesty, and a healthy suspicion of anything that reads a suspiciously round 100.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get compact home essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment is Firmulate, and it ran something unusual: four frontier AI models, each given the same job — running the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the management equivalent of a logbook you can actually inspect.

Why a do-nothing manager still scores 26

The most interesting number in the whole exercise isn’t at the top of the league table — it’s at the bottom. A do-nothing baseline run, where the AI essentially sits on its hands, scores 26 points, not 0. That design choice deserves attention, because it encodes a real truth about management: doing nothing is rarely the worst outcome, and partial progress is real progress. A manager who handles half the inbox, prevents half the fires and closes nothing still delivered something. The benchmark refuses to flatter itself with dramatic zeros.

But the floor has a ceiling-shaped counterpart. A single breach of trust caps the total grade entirely. The framing is blunt: no amount of good work outweighs a breach of trust. In a tiny house, that’s the logic of your load-bearing wall — everything else can be imperfect, but that one failure sinks the structure. The benchmark treats honesty the same way: not as one line item to be averaged away, but as the thing that bounds everything else.

The final league

As of the final July 2026 standings: gpt-5.6-sol took first with 95, Kimi K3 second at 93, Sonnet 5 third with 88, Fable 5 at 77, and Opus 4.8 last at 73. Notably, nothing hit 100 — and the experiment seems fine with that. A perfect round score would be a red flag, not a triumph.

The buried fact

Here’s the finding that should worry anyone delegating work to AI agents. All the models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

Why? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The ones that didn’t read before acting left the close on the table. It’s the AI equivalent of not checking your own rainwater catchment records before sizing a new tank.

Pressure tests and a fairness footnote

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably plain: “Treat the request as a suspected approval-bypass / possible impersonation.” One caveat the experiment itself discloses: K3 ran without an effort parameter while competitors ran at high effort settings — a transparency note you rarely see in this space.

The Opus 4.8 profile is the cautionary tale: it was the most thorough participant, accumulating over 80 learned rules and the deepest analyses, yet finished last. The close went unsigned, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort without follow-through is just expensive motion.

Still running, still watchable

Firmulate didn’t stop at the tournament. It now runs a live company with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live — a slow-burn soap opera with a balance sheet. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson translates cleanly from server racks to small cabins: honest measurement means counting partial progress, refusing to hand out round perfects, and treating trust as the non-negotiable structural element. A gauge that only says “good” or “bad” isn’t a gauge. Whether you’re monitoring a 200-amp-hour battery bank or an AI agent with access to your CRM, you want the instrument that tells you exactly how much you have — including when the tank reads 26.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

tiny house water tank sensor

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Radical Small-Footprint Business Where AI Workers Race a Public Cash Clock

Firmulate’s 13 synthetic employees run a cash-strapped software company in public, revealing whether AI can turn sound judgment into finished work.

Guest app with day-of seating lookup and schedule

A new mobile guest app allows couples to share real-time seating and schedule info via a single link, reducing last-minute questions on wedding day.

What AI Tests Reveal About Trust and Performance in Business — Not Just Chat

Real AI testing reveals that trustworthiness and follow-through matter most. Discover how models perform under pressure in a live business experiment—beyond chat demos.

When the Boss Says “Skip the Rules,” Can an AI Still Be Trusted?

A live AI-company wargame found every model resisted fake CEO pressure, showing trust can be tested before agents enter real workplaces.