
Great Answers Aren’t the Same as Good Management
If you’ve played with an AI chatbot about heat pump sizing or solar payback, you know they write beautifully. Crisp explanations, confident math, polite follow-ups. But homeowners and contractors alike are starting to ask a harder question: when these systems graduate from answering questions to actually running things — a dispatch queue, a service business, a company — what separates a good talker from a good manager?
A live public experiment called Firmulate has been testing exactly that, and the results should make anyone who plans to delegate work to an AI sit up straight.
Same Company, Same Nightmare Week, Four Different AIs
Here’s the setup. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 fielded as participants — were each handed the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome is vibes.
Think of it like this: four crews, each given the same failing boiler to fix in the same freezing house, graded on what actually happened to the pipes — not on how nicely they explained thermodynamics afterward.
The final league table from the July 2026 “Crucible” run:
- gpt-5.6-sol — 95
- Kimi K3 — 93 (with a fairness note: it ran at the API-default effort setting while the others ran at xhigh)
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For perspective, a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.
The Finding That Chat Benchmarks Can’t See
Here’s where it gets interesting for anyone hiring tools, human or otherwise. Every model spotted every crisis. Every model refused every manipulation attempt. And yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Imagine an HVAC estimator who correctly diagnoses the load calculation, writes the perfect proposal, hands it to the homeowner… and never asks for the business. You’d fire them, no matter how smart they sounded.
The buried fact is even more telling. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson translates directly to any trade: the answer is usually in the records — the prior service tickets, the manual, the load history — and the AI that doesn’t read before acting leaves money on the table.
Pressure and Predators
The week also included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant — 80 learned rules added, the deepest analyses in the field — and still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models: effort and insight don’t automatically become follow-through.
It’s Still Running — With Real Money Mechanics
Firmulate isn’t a slide deck. The company is live: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com, and compare models on the benchmarks page. There’s also a genuinely fun twist: 242 real, unedited management decisions power a “guess which model made this call” quiz.
For larger organizations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

As an affiliate, we earn on qualifying purchases.
Management Quality, Not Chat Quality
The takeaway for a home-energy audience is simple. Whether the “company” is a solar installer, a heat pump contractor, or a software shop, the benchmark that matters isn’t how well an AI answers questions — it’s whether it finishes what it starts, reads your files first, stays honest under pressure, and closes the work it earned. A chat demo measures eloquence. A worst-week wargame measures character. If you’re going to put an AI anywhere near your customers, your books, or your crew’s schedule, insist on seeing the second kind of test — because, as the €55k unsigned deal showed, the most expensive failure is the one that never shows up in the transcript.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.