firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In work, as in life, good advice only matters if someone follows through. That familiar gap between knowing and doing is now showing up in a live experiment that asks AI models to run a company through its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate puts frontier AI models in charge of the same small software company, with the same customers, crises and temptations. The decisions are versioned and auditable, and the experiment is watchable as the company runs.

The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. A public cash countdown makes the stakes plain. Each workday is versioned, and the team has built more than 680 self-learned playbook rules. You can watch the experiment at Firmulate.

Amazon

AI task management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the file—and closing the deal

In the final July 2026 Crucible League, gpt-5.6-sol led with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found that buried security weakness, won the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. Its on-record response to a reporter’s “just one yes/no, on background” request was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Knowing is not the same as finishing

The experiment’s central tension is not simply whether a model can diagnose trouble. All could. It is whether the model follows through on the work its own analysis points to. “Same diagnosis, same pitch — no signature” is the gap Firmulate highlights: a model can reach the right conclusion and still leave the valuable action undone.

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped in discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The leaderboard suggests that detail and eloquence alone do not settle who is ready to take responsibility for a business task.

Make the test your own

For anyone weighing AI tools for everyday work, the lesson is practical: a polished answer is only part of the job. Does the model read the relevant material, protect trust under pressure and complete the action? Those are questions a chat demonstration may not answer. Firmulate offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice.

Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. The complete results and plain-language findings are at Firmulate’s benchmark page.

Fairness note: K3 ran without an effort parameter (API default), while the other models ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The takeaway

Kimi K3’s second-place finish shows the league is open, while the missed deal among capable models shows why a familiar name or strong demo is not enough. Before choosing an AI to handle real work, test whether it can turn sound judgment into trustworthy follow-through.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Discover Heartfelt Sad Emotional Quotes for Healing

Find solace in heartfelt sad emotional quotes that resonate with your feelings. Discover words of comfort and healing to help you through difficult times and process your emotions.

Inspirational Quotes on the Importance of Fathers

Tune into heartfelt quotes celebrating the transformative impact of fathers, leaving you inspired to discover the profound influence of paternal love and guidance.

Happy Birthday to Daughter From Father: Celebrate Her Day

A father's heartfelt birthday wishes for his daughter will make her day truly special – discover the emotional connection in this heartfelt celebration.

Discover Fresh Quotes New: Inspiring Words for You

Explore a treasure trove of quotes new to inspire and uplift you. Find fresh perspectives, motivational words, and thought-provoking insights for every occasion.