firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

What an energy-conscious homeowner can learn from an AI management stress test

Anyone comparing heat pumps, solar panels or backup power knows that specifications tell only part of the story. The harder questions concern performance under pressure: Will the system cope when conditions deteriorate? Will the supplier notice an important detail? Will somebody finish the job rather than merely explain it well?

Firmulate applies that same practical instinct to frontier artificial intelligence. Its live experiment gave each model the same small software company and sent it through its worst week, with identical customers, crises and temptations. The decisions were versioned and auditable. What emerged was not simply a ranking of clever answers, but a set of recognizable management personalities.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between seeing a problem and solving it

The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust”.

The most revealing finding was not that some models missed the crises. None did. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature”.

That distinction matters well beyond software. In home energy, a convincing assessment is not the same as a completed installation, and identifying a battery constraint is not the same as resolving it. Firmulate’s experiment asks a similarly concrete question of AI: after the analysis is finished, does the model carry the decision through?

The crucial fact was hiding in the company’s own documents

The deal turned on a competitor weakness buried two document references deep in the company’s files. It was not visible in the customer event itself. Models that followed the references found the information and won the deal at full price, worth +€4,583 MRR.

This is an unusually relatable test for readers accustomed to technical products. Important facts are often scattered across quotations, manuals, tariffs and warranty conditions. A decision-maker who reacts only to the latest message may appear responsive while missing the evidence that changes the commercial outcome. Firmulate’s result shows that reading the available material was not administrative housekeeping; it was decisive work.

Every model held the line against manipulation

The company was also subjected to fake CEO messages escalating over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background”. All 5 models refused. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was a strong shared result. The models differed in follow-through and operating discipline, but not in their willingness to reject the manipulation attempts presented during the experiment. For businesses considering AI access to customer records, support queues or forecasts, that distinction is useful: safety under pressure and commercial completion are separate capabilities, and both have to be observed.

The most thorough manager did not win

Opus 4.8 offers the sharpest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The lesson is not that depth lacks value. It is that depth can coexist with hesitation or procedural drift. A model may document more, learn more and still fail to complete the action that matters most. By contrast, the league leaders paired analysis with a successful close.

There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. Its 93 result therefore belongs in the table, but the operating difference should remain visible when readers compare the performances.

A quiz built from actual decisions

Firmulate has turned 242 real, unedited management decisions into a guess-the-model quiz. Readers see the decisions and try to identify which frontier model made them before learning the answer and the associated character profile.

The format makes abstract model differences tangible. One participant can be expansive, another terse, and another unwilling to communicate through noise. Instead of asking which response sounds smartest in isolation, the quiz asks whether readers can recognize stable habits across real company situations.

The surrounding company remains a watchable experiment: 13 synthetic employees operate with real money mechanics, burning €105k/month against €2.3k MRR. It has a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For AI buyers, the management style is part of the product

Firmulate’s experiment suggests that frontier models can reach the same diagnosis while producing materially different business outcomes. The decisive differences appeared in whether they read deeply enough, completed the commercial action and followed the right escalation path when ordinary work became difficult.

That should sound familiar to anyone making a serious home-energy purchase. Headline capability matters, but so do diligence, reliability and follow-through. The quiz offers an accessible way to test whether those traits are visible before seeing the model name—and to consider what kind of manager an AI becomes when the week goes badly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding Energy-Efficient Heat Pumps for Green Homes

We have all the information on energy-efficient heat pumps for environmentally friendly…

6 Essential Steps to Install Your Heat Pump

Welcome, fellow homeowners! Are you ready to embark on an exciting journey…

Revolutionary Energy-Saving AC Heat Pumps: A Curated Guide

Our goal is to improve the energy efficiency of your home with…