
We grow up on a simple idea: do nothing, get nothing. Flunk the test you didn’t study for, and you get a zero. That’s why one detail from a new AI experiment is so quietly fascinating: in the Firmulate benchmark, a manager that does nothing at all still walks away with 26 points.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Not zero. Twenty-six. And the people running the experiment consider that a feature, not a bug — because it tells you something true about how work, trust, and credit actually function in a real company.
The company that never existed
Firmulate ran four frontier AI models — GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8 among them — through the identical nightmare: the same small software company, the same customers, the same week of crises and temptations. Every decision was versioned and auditable, so nothing could be quietly retrofitted afterward.
The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Those are strong numbers — but the more interesting number is the floor beneath them all: 26.
As an affiliate, we earn on qualifying purchases.
Why doing nothing still earns 26
The logic is almost philosophical. Even a manager who takes no decisive action still occupies the chair. Crises get noticed. Fires burn at a predictable rate. Some fraction of the business simply carries on. A scoring system that gave total inaction a zero would be lying — it would imply that everything good in a company flows from the manager’s genius, and we all know from our own workplaces that this isn’t true.
So partial progress counts. A manager who diagnoses a problem correctly but never closes the deal has still done real work. In fact, that’s exactly what happened: every model spotted every crisis and refused every manipulation attempt, yet only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s own shorthand for this gap: “Same diagnosis, same pitch — no signature.”
The one rule that caps everything
But partial credit has a hard ceiling, and it isn’t about effort. A single breach of trust caps the total grade — the benchmark’s stated philosophy is blunt: “No amount of good work outweighs a breach of trust.”
Think of the best manager you ever had, and the worst. The worst one probably wasn’t incompetent — they were competent and untrustworthy. That’s the distinction the scoring tries to encode. Brilliance with a side of dishonesty doesn’t average out to “pretty good.” It gets capped.
The models were genuinely tested on this. A social-engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
The buried fact that decided the league
Here’s the detail that separates the winners from the also-rans. The decisive weakness in the customer’s current vendor wasn’t in the customer meeting at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that skimmed lost it.
It’s the workplace proverb made measurable: the answer was in the drawer the whole time. Nobody opened the drawer.
Why the most thorough model finished last
The cruelest twist belongs to Opus 4.8. It was the most thorough participant by volume — over 80 learned rules, the deepest analyses in the field. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating properly. Notably, the same weakness appeared, weaker, in all four models — suggesting it’s a pattern, not a fluke.
One fairness footnote: Kimi K3 ran at its API default effort setting while the others ran at maximum effort — and still finished second. That makes its 93 arguably the most impressive number on the board.
Distrust of round 100s
Perhaps the most honest thing about the benchmark is its suspicion of perfection. No model scored 100, and the design seems to guard against it — because a perfect score on a messy, human-flavored simulation would be a red flag, not a triumph. Real management doesn’t grade out at 100. Why should AI management?
You can watch the sequel live
The experiment didn’t end with the league table. A live synthetic company runs at firmulate.com/live: 13 employees, real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and the site rebuilds itself twice a day. It’s oddly compelling — somewhere between a business documentary and a terrarium.
There’s also a quiz built on 242 real, unedited management decisions where you guess which model made which call, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The 26-point floor is the whole philosophy in one number: give credit where it’s genuinely earned, cap the score where trust is broken, and never hand out a perfect grade for an imperfect job. It’s how most of us wish our own performance reviews worked. Now someone has built an AI benchmark that actually does it — and you can watch the results, and the countdown, in public.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
