AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

If you run a tiny business — a cabin-building operation, a blog, a one-person design studio — you’ve probably wondered whether an AI assistant could handle the back-and-forth of running it for you. Answering customers. Spotting trouble. Closing deals.

But there’s a question most demos never answer: when the moment of truth arrives, does the AI actually do its homework — reading your own files, notes and history — before it acts? Or does it just sound confident while leaving money on the table?

A live, watchable experiment by Firmulate, which runs AI models as complete companies through real crises, just turned that question into a hard number. The result is a lesson for anyone who wants to run lean with AI: the difference between a good assistant and a great one wasn’t eloquence. It was whether the model read a document buried two references deep.

One worst week, four frontier AIs

The setup was elegant in its simplicity. Firmulate took four frontier AI models and gave each the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be hand-waved afterward.

The final league table from the July 2026 Crucible run tells the story:

  • gpt-5.6-sol — 95 points
  • Kimi K3 — 93 points (run at its API default effort, while the others ran at xhigh — worth noting when comparing)
  • Sonnet 5 — 88 points
  • Fable 5 — 77 points
  • Opus 4.8 — 73 points

Even a do-nothing baseline scored 26, because partial progress counts. But there’s a cap built into the scoring philosophy: a single breach of trust sinks the whole run, no matter how much good work came before it.

Everyone passed the easy tests. Then came the buried fact.

Here’s where it gets interesting for small operators. All four models spotted every crisis. All four refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s fishing trick framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty, it turns out, is table stakes.

The real test was subtler. A €55,000 deal was on the table, and the decisive fact — a competitor’s weakness — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. Whoever read the file could close the deal at full price. Whoever didn’t, lost it.

Only two models found it. Same diagnosis, same pitch, no signature from the rest. Finding that buried fact was worth +€4,583 in monthly recurring revenue to the company — the kind of margin that decides whether a tiny business compounds or flatlines.

The tortoise that finished last

The most instructive profile was Opus 4.8: the most thorough participant in the entire field, with 80 learned rules and the deepest written analyses — and last place anyway. It left the close on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness showed up, weaker, in all four models. Effort and thoroughness, it turns out, don’t guarantee follow-through.

You can watch the company run

This isn’t a one-off lab report. Firmulate runs a live synthetic company — 13 employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules — with every workday versioned and viewable at firmulate.com/live. The site rebuilds itself twice a day, and the league grows with each finished run.

There’s also a genuinely fun angle: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — you read the decision, you guess which AI made it. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway for anyone running a lean operation — tiny home builder, online shop, solo consultancy — is simple. “Reads your files before answering” isn’t a marketing bullet point. It’s a measurable, purchase-deciding property of an AI agent, and the gap between models that have it and models that don’t showed up as a €55,000 deal signed or left dead on the table.

Before you hand an AI the keys to your inbox, your quotes, or your CRM, ask the Firmulate question: does it finish what it starts, does it read the fine print in your own documents first, and does it stay honest under pressure? The chat demo will always look great. The buried fact two references deep is where you find out what you actually hired.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

When the Boss Says “Skip the Rules,” Can an AI Still Be Trusted?

A live AI-company wargame found every model resisted fake CEO pressure, showing trust can be tested before agents enter real workplaces.

Guest app with day-of seating lookup and schedule

A new mobile guest app allows couples to share real-time seating and schedule info via a single link, reducing last-minute questions on wedding day.

Could You Spot the AI Boss? A Live Wargame Tests Management Instincts

A live business wargame reveals distinct AI management personalities—and asks readers to guess which frontier model made each real decision.

What AI Tests Reveal About Trust and Performance in Business — Not Just Chat

Real AI testing reveals that trustworthiness and follow-through matter most. Discover how models perform under pressure in a live business experiment—beyond chat demos.