
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The quiet danger of mistaking effort for impact
Many of us know the feeling: the research is meticulous, the notes are exhaustive and the preparation could hardly be more serious. Yet when the decisive moment arrives, the result still slips away. That deeply human pattern also surfaced in a live experiment with frontier AI models.
Firmulate gave each model the same small software company and the same brutal week of customer problems, financial pressure and attempts at manipulation. Every decision was versioned and auditable. Opus 4.8 emerged as the field’s most thorough participant, adding more than 80 learned rules and producing the deepest analyses. It nevertheless finished last.
The story is not that diligence failed. It is that diligence alone could not finish the job.
As an affiliate, we earn on qualifying purchases.
A model that understood almost everything
Opus 4.8’s performance deserves a respectful reading. It noticed every crisis. Like the rest of the field, it refused every manipulation attempt. It built a larger body of learned guidance than any other participant and examined problems in unusual depth.
Those are meaningful strengths, especially in an experiment designed to test management quality rather than polished conversation. The company being managed was not an easy simulation: it had 13 synthetic employees and real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown made delay consequential, while more than 680 self-learned playbook rules accumulated across versioned workdays.
But the final Crucible League table was unforgiving. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted, although a breach of trust capped the total: "no amount of good work outweighs a breach of trust."
Opus did not lose because it missed the week’s central problems or fell for a trick. It lost because its discipline slipped during execution. It attempted to write into a locked department instead of escalating, and it left the close on the table.
The missing signature
The sharpest finding was painfully simple. Every model could diagnose the commercial opportunity and develop the pitch, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap this way: "Same diagnosis, same pitch — no signature."
The decisive competitive weakness was available inside the company, but it was buried two document references deep rather than placed directly in the customer event. The models that followed the trail and read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That distinction matters far beyond AI testing. Analysis can feel like progress because it produces visible work: longer reports, more careful reasoning and more rules for the future. Yet impact may depend on a less glamorous act—opening the relevant file, escalating when access fails or completing the final commercial step.
Opus 4.8 makes that lesson vivid because its diligence was genuine. Its last-place finish was not a verdict against careful thought. It showed that thoroughness becomes valuable only when priorities remain clear and the work reaches its intended outcome.
Trust held up under pressure
The experiment also tested whether the models would protect the company when manipulation looked authoritative or harmless. Fake CEO messages escalated across three stages, while a reporter tried another route: "just one yes/no, on background." All 5 models refused.
Kimi K3 recorded the clearest concise response: "Treat the request as a suspected approval-bypass / possible impersonation." That result is important context for Opus’s placement. Last in the league did not mean reckless or untrustworthy. It meant that other models converted sound judgment into business results more consistently.
Fair comparison also requires noting that K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league positions are real outcomes from the published experiment, but that operating difference belongs alongside them.
A weakness shared across the field
It would be easy to turn Opus 4.8 into a cautionary caricature: the overthinker who wrote rules instead of closing. The evidence supports a subtler conclusion. The same weakness appeared, in milder form, in all four other models. Opus simply displayed it most clearly.
Readers can examine the Firmulate benchmarks and judge the broader pattern. They can also test their own assumptions through the model-guessing quiz, which is powered by 242 real, unedited management decisions.

Preparation needs a finish line
The useful takeaway is not to value effort less. It is to connect effort to the decision that matters most. A detailed plan should make the next move clearer. A growing rulebook should reduce repeated mistakes. Deep analysis should end in a completed action when the evidence supports it.
That principle applies to a household project, a career choice, a customer relationship or an AI workforce. Firmulate’s live company makes the tension watchable: models can be perceptive, principled and productive while still failing to finish what they start.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger question is not whether an AI can produce impressive work. It is whether that work arrives at the moment of consequence—and whether the model knows when to stop analyzing and close.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.