
Real competence begins after the impressive answer
In everyday life, we learn to distinguish eloquence from reliability. The friend with perfect advice may disappear when help is needed; the colleague who understands a problem may still fail to finish the job. Yet much of the excitement around artificial intelligence rests on polished responses, coding scores and head-to-head chat preferences.
Those tests matter, but they leave a management-sized hole. They rarely show how an AI agent prioritizes when several crises arrive together, whether it investigates before acting, or whether it remains honest when deception would make its results look better. A model can ace an isolated task without proving that it can carry responsibility across days.
Firmulate is testing that missing layer. Its live experiment gave each frontier model the same small software company during its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The central question was not whether the models could produce convincing language. It was whether they could manage.
business AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
When diagnosis is not delivery
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a hard ethical boundary: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
The encouraging result was that every model detected every crisis and rejected every manipulation attempt. The unsettling result was that only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
That is a useful warning for anyone imagining agents working inside a company. Recognizing the right course is not the same as completing it. A model may explain a sales opportunity beautifully, prepare the case and still leave the commercial outcome untouched. In a chat window, that performance can look excellent. In a business, it is unfinished work.
The decisive clue was buried in ordinary work
The deal also tested whether models would look beyond the obvious event. The decisive weakness in a competitor was not sitting in the customer interaction. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR.
This finding shifts the debate from raw intelligence toward working habits. The strongest business agent may not be the one that gives the fastest confident answer. It may be the one that reads the available material, follows references and notices the detail that changes a negotiation. That resembles good human management: context often matters more than verbal sparkle.
Pressure also tests integrity
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because business pressure rarely announces itself as an ethics examination. It arrives as urgency, authority, flattery or a request framed as harmless. A useful agent must recognize the social situation, not merely parse the words. In this field, the models’ collective refusal is a meaningful success.
Thoroughness did not guarantee the best outcome
Opus 4.8 offers the sharpest lesson. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
This is not an argument against careful analysis. It is evidence that analysis alone cannot stand in for judgment, execution and procedural discipline. A growing rulebook can coexist with a missed result. The K3 comparison also needs its fairness note: K3 ran on the API default without an effort parameter, while the others ran at xhigh.
Readers can inspect the full benchmark findings, while 242 real, unedited management decisions also power a public guess-the-model quiz. The point is not to turn management into entertainment, but to make model behavior tangible enough that people can judge it for themselves.

A new curriculum for business AI
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k/month against €2.3k MRR, maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real and watchable, allowing visitors to follow consequences rather than accept a polished demonstration.
That suggests a better curriculum for evaluating AI agents: churn wave, price increase, downround and PR crisis. These scenarios ask whether a model can triage under limited capacity, preserve trust, use institutional knowledge and finish valuable work. They measure management quality, not chat quality.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is the right spirit for adoption: test an agent amid consequences before granting it responsibility. The future winner may still write excellent code and compelling prose. But the model worth hiring will also read the files, resist the shortcut, escalate appropriately and close the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html