firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Buy a Heat Pump on the Brochure Alone

Nobody in the home-energy world buys a heat pump based on how nicely the manufacturer chats about COP ratings. You want the real-world numbers: how it performs on the coldest day, under load, over a whole season. So why do companies pick AI models — systems that will touch customer data, support queues, and money — based on a chat demo? A live experiment at Firmulate just showed why that’s a bet, not a decision.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five AI Models, One Terrible Week

Firmulate ran a public wargame: five frontier AI models each ran the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. The company is real running software with 13 synthetic employees, real money mechanics (burning €105k/month against €2.3k MRR), a public cash countdown, and 680+ self-learned playbook rules. Every decision is versioned and auditable, and you can watch it at firmulate.com/live.

Amazon

enterprise AI evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Final League Table

As of the final July 2026 standings, Moonshot’s Kimi K3 — the newcomer — finished second with 93, ahead of Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol scored higher, at 95. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

Amazon

AI decision validation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What K3 Actually Did

  • Found the buried security needle — the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the €55,000 deal at full price, worth +€4,583 in MRR.
  • Signed the €55,000 deal its own analysis had earned.
  • Saved a churning customer.
  • Resisted all three manipulation baits — fake CEO messages escalating over three stages, plus a reporter trick (“just one yes/no, on background”). All five models refused, but K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • Recorded only one deviation all week — the cleanest discipline in the field.
Amazon

AI trustworthiness testing kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Should Worry Every Buyer

Here’s the part no chat demo reveals: all five models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The gap between “understands the problem” and “finishes the job” is invisible in a demo.

Then there’s Opus 4.8: the most thorough participant, with 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. A weaker version of that same weakness showed up in all four other models.

Try It Yourself

Firmulate’s quiz lets you guess which model made which call, powered by 242 real, unedited management decisions (firmulate.com/quiz.html). Full benchmarks are at firmulate.com/benchmarks, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Lesson for Anyone Buying Smart Systems

Whether it’s a heat pump or an AI agent, the specs lie less than the demo but more than the season. The league is open now — a newcomer from Moonshot beat three of four Western frontier models. If AI will touch your business, picking a model without running your own test is no longer a procurement choice; it’s a gamble.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Seasonal Performance of Heat Pumps in Cooling Mode

Ineffective cooling performance can be improved by understanding the factors that influence heat pump seasonal efficiency. Continue reading to discover how.

I Created A Pretty IKEA BILLY Hack That Seriously Cuts Clutter

A DIY enthusiast has shared a creative hack for the IKEA BILLY bookcase that effectively minimizes clutter, gaining rapid attention online.

Boosting Heat Pump Efficiency: A Sustainable Building Guide

AIThis post was created with the assistance of artificial intelligence (AI). Were…

Top 5 Steps to Install a Heat Pump

AIThis post was created with the assistance of artificial intelligence (AI). Did…