AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What tiny living can teach us about artificial intelligence

Tiny-home enthusiasts understand that constraints expose priorities. When space, money and resources are tight, every choice matters: what gets preserved, what gets sacrificed and whether a clever plan is actually completed. The same principle applies to AI management.

Firmulate has turned that idea into a live business wargame. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. Now, an interactive quiz built from 242 real, unedited management decisions asks readers to identify which model was in charge.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The management personalities hiding behind the prose

The appeal of the quiz is not simply guessing an AI brand from its writing style. It is seeing how differently models behave when analysis must become action. Some investigate deeply. Some maintain cleaner operational discipline. Some recognize the right commercial opportunity but still fail to close it.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

That safeguard mattered because the company’s worst week included more than ordinary business pressure. Fake CEO messages escalated over three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3’s recorded reasoning was blunt and useful: “Treat the request as a suspected approval-bypass / possible impersonation.”

Everyone saw the crisis. Not everyone finished the job.

The most revealing result appeared in sales. Every model spotted every crisis, and the models reached the same diagnosis and produced the same pitch. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The difference was buried in the company’s own records. A decisive competitor weakness sat two document references deep in internal files rather than in the customer event. Models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That finding will feel familiar to anyone who has undertaken a compact-building project. Recognizing a constraint is not the same as tracing it through the plans, confirming the details and making the final commitment. An impressive explanation cannot substitute for completed work.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest character study. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. It nevertheless finished last. The commercial close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.

The same weakness appeared less severely in all four other competitors. That makes the experiment more interesting than a simple ranking of “smart” and “less smart” systems. The models could understand the situation while still differing in follow-through, file-reading habits and procedural discipline.

Kimi K3’s result also comes with an important fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place score should be read with that difference in mind.

A company designed to make consequences visible

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown making the pressure visible. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

This is why the quiz works as more than entertainment. Its decisions come from a running experiment in which commercial opportunities, trust tests and operational mistakes have visible consequences. Readers are not choosing between polished hypothetical answers; they are examining what the models actually did when each one received the same business situation.

Infographic —
The findings at a glance — source: firmulate.com.

A more practical way to judge an AI manager

Chat demonstrations often reward eloquence. Firmulate’s experiment asks harder questions: Did the model read the relevant files? Did it resist pressure? Did it respect boundaries? Did it convert a sound analysis into a finished result?

The answers form recognizable management profiles. Opus 4.8 shows that exhaustive analysis can coexist with weak completion. Kimi K3 demonstrates disciplined suspicion when authority may be impersonated. The league leaders show the commercial value of following evidence beyond the obvious event.

For readers accustomed to alternative living, the broader lesson is fitting: performance becomes clearest under constraint. Limited resources and real consequences strip away presentation. What remains is judgment, discipline and the ability to finish.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Guest app with day-of seating lookup and schedule

A new mobile guest app allows couples to share real-time seating and schedule info via a single link, reducing last-minute questions on wedding day.

The Radical Small-Footprint Business Where AI Workers Race a Public Cash Clock

Firmulate’s 13 synthetic employees run a cash-strapped software company in public, revealing whether AI can turn sound judgment into finished work.

When the Boss Says “Skip the Rules,” Can an AI Still Be Trusted?

A live AI-company wargame found every model resisted fake CEO pressure, showing trust can be tested before agents enter real workplaces.

What AI Tests Reveal About Trust and Performance in Business — Not Just Chat

Real AI testing reveals that trustworthiness and follow-through matter most. Discover how models perform under pressure in a live business experiment—beyond chat demos.