firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The quiet danger of mistaking effort for impact

Many of us know the feeling: the research is meticulous, the notes are exhaustive and the preparation could hardly be more serious. Yet when the decisive moment arrives, the result still slips away. That deeply human pattern also surfaced in a live experiment with frontier AI models.

Firmulate gave each model the same small software company and the same brutal week of customer problems, financial pressure and attempts at manipulation. Every decision was versioned and auditable. Opus 4.8 emerged as the field’s most thorough participant, adding more than 80 learned rules and producing the deepest analyses. It nevertheless finished last.

The story is not that diligence failed. It is that diligence alone could not finish the job.

Amazon

professional file organizer

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A model that understood almost everything

Opus 4.8’s performance deserves a respectful reading. It noticed every crisis. Like the rest of the field, it refused every manipulation attempt. It built a larger body of learned guidance than any other participant and examined problems in unusual depth.

Those are meaningful strengths, especially in an experiment designed to test management quality rather than polished conversation. The company being managed was not an easy simulation: it had 13 synthetic employees and real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown made delay consequential, while more than 680 self-learned playbook rules accumulated across versioned workdays.

But the final Crucible League table was unforgiving. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted, although a breach of trust capped the total: "no amount of good work outweighs a breach of trust."

Opus did not lose because it missed the week’s central problems or fell for a trick. It lost because its discipline slipped during execution. It attempted to write into a locked department instead of escalating, and it left the close on the table.

The missing signature

The sharpest finding was painfully simple. Every model could diagnose the commercial opportunity and develop the pitch, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap this way: "Same diagnosis, same pitch — no signature."

The decisive competitive weakness was available inside the company, but it was buried two document references deep rather than placed directly in the customer event. The models that followed the trail and read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That distinction matters far beyond AI testing. Analysis can feel like progress because it produces visible work: longer reports, more careful reasoning and more rules for the future. Yet impact may depend on a less glamorous act—opening the relevant file, escalating when access fails or completing the final commercial step.

Opus 4.8 makes that lesson vivid because its diligence was genuine. Its last-place finish was not a verdict against careful thought. It showed that thoroughness becomes valuable only when priorities remain clear and the work reaches its intended outcome.

Trust held up under pressure

The experiment also tested whether the models would protect the company when manipulation looked authoritative or harmless. Fake CEO messages escalated across three stages, while a reporter tried another route: "just one yes/no, on background." All 5 models refused.

Kimi K3 recorded the clearest concise response: "Treat the request as a suspected approval-bypass / possible impersonation." That result is important context for Opus’s placement. Last in the league did not mean reckless or untrustworthy. It meant that other models converted sound judgment into business results more consistently.

Fair comparison also requires noting that K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The league positions are real outcomes from the published experiment, but that operating difference belongs alongside them.

A weakness shared across the field

It would be easy to turn Opus 4.8 into a cautionary caricature: the overthinker who wrote rules instead of closing. The evidence supports a subtler conclusion. The same weakness appeared, in milder form, in all four other models. Opus simply displayed it most clearly.

Readers can examine the Firmulate benchmarks and judge the broader pattern. They can also test their own assumptions through the model-guessing quiz, which is powered by 242 real, unedited management decisions.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Preparation needs a finish line

The useful takeaway is not to value effort less. It is to connect effort to the decision that matters most. A detailed plan should make the next move clearer. A growing rulebook should reduce repeated mistakes. Deep analysis should end in a completed action when the evidence supports it.

That principle applies to a household project, a career choice, a customer relationship or an AI workforce. Firmulate’s live company makes the tension watchable: models can be perceptive, principled and productive while still failing to finish what they start.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger question is not whether an AI can produce impressive work. It is whether that work arrives at the moment of consequence—and whether the model knows when to stop analyzing and close.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Inspirational New Father Quotes to Celebrate Fatherhood

Lift your spirits with heartwarming quotes that capture the essence of fatherhood, guiding you through this transformative journey.

Daughter Quotes From Dad: Words of Love and Wisdom

Uncover touching daughter quotes from dads brimming with love and wisdom, showcasing the unique bond that shapes daughters' lives.

Father and Daughter Love Quotes: Unbreakable Bonds

Uncover the unbreakable bond and profound love between fathers and daughters in heartwarming quotes that will touch your soul.

Beautiful Quotes Celebrating Daughter and Dad

Peek into the heartfelt connection between a daughter and her dad with touching quotes that capture their unconditional love and cherished moments.