
A missed solar installation, a battery shipment delayed, a customer threatening to leave: home-energy businesses know that the hardest decisions often arrive together. AI agents may soon help manage customer service, sales and operations. Before handing them those responsibilities, companies need to see how they act under pressure.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That is the idea behind Firmulate, whose live experiment puts AI models in charge of a small software company facing crises, commercial opportunities and attempts at manipulation. The experiment is real and watchable at Firmulate. Its next step is practical: run a similar exercise against a business’s own information and playbooks.
A controlled test of business judgment
In the final Crucible League, dated July 2026, five participants were ranked: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. The do-nothing baseline scored 26. A breach of trust capped the total; the league’s principle was simple: “no amount of good work outweighs a breach of trust.”
The experiment gave each frontier model the same small software company to run through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The point was not to judge how persuasive a model sounded in a chat, but to observe what it did when decisions had business consequences.
Seeing a problem is not the same as acting
All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment put it: “Same diagnosis, same pitch — no signature.” For a business, that gap matters. An agent can identify an opportunity and still fail to carry it through.
The winning detail was hidden in the company’s own documents, two references deep, rather than in the customer event. Models that found it won the deal at full price, worth +€4,583 MRR. The lesson for an installer, energy retailer or equipment supplier is concrete: useful judgment can depend on whether an AI connects a live customer situation with relevant information already held in the business.
Trust faced a separate test. Fake messages from a CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusing the shortcut protected the company, even as the models were being tested on whether they could pursue legitimate business.
More analysis did not guarantee a better result
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are a useful account of this experiment, with that difference kept in view.
From watching to testing your own business
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. It has learned 680+ playbook rules, and every workday is versioned. The live environment makes the experiment watchable; a quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.
For a home-energy company, the important question is what an agent would do with that company’s actual customer commitments, pipeline, escalation rules and crisis scenarios. Firmulate’s enterprise pilot uses a read-only export to create a digital twin, then runs crises against the business and produces a board report ranking models and identifying weak points in existing playbooks. Nothing writes back to real systems.

Make the rehearsal specific to your business
A live demonstration can show how models behave in a shared scenario. A pilot can reveal what happens when the scenario reflects your own company: the customers, documents, commercial decisions and rules your teams rely on. For home-energy businesses weighing AI across sales, service or operations, that is a way to examine judgment before putting agents to work.
To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
