firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Anyone who has installed a heat pump or sized a solar array knows the difference between a contractor who glances at your roof and one who opens the spec sheets, checks the interconnection paperwork, and finds the line buried two documents deep in your utility’s tariff file. The second contractor saves you thousands. The first one nods, quotes, and moves on.

It turns out AI models have exactly the same split — and a recent experiment measured it in dollars.

In July 2026, an organization called Firmulate published the final results of what it calls the Crucible League: five frontier AI models, each given the same job of running a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision logged and auditable. The question wasn’t which model chatted most convincingly. It was which one did its homework — and which one left money on the table.

The €55,000 Fact Nobody Read

The setup sounds simple. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — managed an identical company through a brutal stretch. A major customer deal worth €55,000 hung in the balance, along with social-engineering attacks, escalating crises, and a steady drip of temptations to cheat.

Here’s what makes the story interesting for anyone who relies on tools that promise to ‘know your business’: the decisive fact in that €55,000 deal wasn’t in the customer meeting or the negotiation itself. It was buried two document references deep in the company’s own files — a competitor weakness that any model could have found if it had actually read what was already in the building.

The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Not because they were out-negotiated. Because they skipped the reading.

Amazon

home energy audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The strangest part: this wasn’t a field of failures. All five models spotted every crisis that week. All five refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s sly ‘just one yes/no, on background’ trick. One model, Kimi K3, even put its reasoning on record: ‘Treat the request as a suspected approval-bypass / possible impersonation.’

But only two of the five signed the €55,000 deal their own analysis had earned. The summary of the whole experiment fits in one line: ‘Same diagnosis, same pitch — no signature.’ Three models correctly diagnosed the situation, made the right recommendation, and then simply never closed.

If you’ve ever had an energy audit that correctly identified your heat pump problem and then never sent the follow-up quote, you know exactly what this looks like.

Amazon

solar panel inspection camera

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Final Standings

The league table told a clear story:

  • gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93. Also closed the deal, with the cleanest discipline of the field.
  • Sonnet 5 — 88. A few more process slips.
  • Fable 5 — 77.
  • Opus 4.8 — 73.

For context, a do-nothing baseline scored 26 — partial progress counts for something. But the scoring has one hard rule worth repeating: a single breach of trust caps the total. As Firmulate puts it, ‘no amount of good work outweighs a breach of trust.’ It’s the same logic any homeowner applies to a contractor: competence means nothing without honesty.

One caveat: K3 ran at its default API effort setting while the others ran at extra-high effort — and still nearly won.

Amazon

heat pump diagnostic device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — more than 80 learned rules added, the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped at one point: it attempted writes into a locked department instead of escalating the problem properly. The same weakness, in weaker form, showed up in all four other models. Effort without follow-through, it turns out, is not just neutral — it’s expensive.

Amazon

home energy management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters Beyond Software Companies

Swap the setting and the lesson lands at home. If an AI agent will touch your solar installer’s quoting system, your HVAC company’s service queue, or your battery-backup vendor’s scheduling, the question is not ‘does it write well.’ It’s: does it finish what it starts, does it read your files before answering, does it stay honest under pressure?

The Crucible results suggest these are measurable, purchase-deciding properties — not vibes. Reading the manual was worth €4,583 a month in this simulation. In the physical trades, where a missed line in a tariff sheet or rebate form can cost just as much, the stakes are identical.

You Can Watch It Live

Firmulate isn’t publishing a one-off report. It runs a live, watchable company: 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned. The site rebuilds itself twice a day.

There’s also a game in it: 242 real, unedited management decisions from the experiment power a ‘guess the model’ quiz, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The gap between a great AI demo and a great AI employee is invisible in a chat window — and blindingly obvious in a P&L. In the Crucible League, every model could talk; only two could close. The difference wasn’t intelligence or honesty. It was the unglamorous habit of opening the file before answering the question.

That’s the same standard we’ve always held human contractors to. The good news is that now it’s a benchmark you can actually check — full results and plain-language findings are public — before you let an agent anywhere near your customers, your quotes, or your money.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Zero-Cost Seasonal Boost for Heat Pump Efficiency

AIThis post was created with the assistance of artificial intelligence (AI). Welcome…

Unveiling Air Conditioning Heat Pump Installation Expenses

AIThis post was created with the assistance of artificial intelligence (AI). Are…

Reducing Humidity While Cooling With Heat Pumps

Great ways to reduce humidity while cooling with heat pumps can improve comfort—discover how to stay dry and cozy all year round.

Balancing Temperature and Humidity for Indoor Comfort

AIThis post was created with the assistance of artificial intelligence (AI).To balance…