firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A workplace lesson hiding inside an AI wargame

In everyday life, reliability often comes down to an unglamorous habit: doing the reading before making a decision. The same appears to be true for artificial intelligence. An agent may recognize a problem, write a persuasive response and sound thoroughly prepared. But if it fails to follow the trail through the available documents, it can still miss the fact that determines whether the work succeeds.

Firmulate turned that distinction into a measurable test. Its live experiment placed frontier AI models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable. The decisive challenge was not merely identifying an opportunity. It was finding a competitor weakness buried two document references deep in the company’s own files—and then using it to close a €55,000 deal.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between sounding ready and being ready

Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

That result exposes a weakness that ordinary chatbot demonstrations rarely reveal. A model can produce the right diagnosis and even recommend the right commercial approach while failing to complete the action that gives the work value. In this case, the buried fact was not present in the customer event. The models had to consult the company’s files, follow two references and recognize why the discovered weakness mattered. Those that did won the deal at full price, worth +€4,583 MRR. Those that did not lost it automatically.

This makes “reads your files before answering” more than a reassuring product claim. In Firmulate’s experiment, it became a purchase-deciding property. The models were not separated by whether they could discuss the opportunity intelligently. They were separated by whether they gathered the evidence required to finish the job.

A demanding company, not a tidy prompt

The setting is designed to resemble an operating business rather than an isolated question. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing observers to inspect what happened rather than relying on a polished retrospective.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.” The complete standings and findings are available on the Firmulate benchmarks page.

The ranking also carries an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That distinction matters when readers compare performances, even though the central outcome remains straightforward: K3 was among the two models that found what mattered and completed the deal.

Reading depth did not guarantee overall discipline

Opus 4.8 offers the most revealing cautionary profile. It was the most thorough participant, added 80 learned rules and produced the deepest analyses, yet finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This is a useful corrective to the idea that more analysis inevitably produces better management. Thoroughness can help an agent uncover context, but practical performance also requires judgment, follow-through and respect for boundaries. The winning behavior was not simply reading more. It was reading the right material, applying the discovered fact and carrying the decision through without losing operational discipline.

Pressure tested without compromising trust

The models also faced fake CEO messages that escalated over three stages, plus a reporter attempting to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result matters because useful workplace agents must do two things at once: pursue legitimate goals and resist shortcuts that violate trust. Firmulate’s test suggests those abilities can coexist. The harder distinction emerged elsewhere—in whether the agent could convert available information into a completed business outcome.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What buyers should ask before hiring an AI agent

A polished answer is not enough evidence that an agent will perform reliably inside a business. Buyers should ask whether it examines relevant files, follows references, finishes actions and escalates when it encounters a boundary. Firmulate makes those behaviors visible through a live, watchable company rather than a scripted chat showcase.

Readers can also test their own instincts through Firmulate’s “guess the model” quiz, powered by 242 real, unedited management decisions. Enterprises can go further by running the same wargame against a read-only export of their own business; nothing writes back to real systems.

The broader lesson is simple and surprisingly human: preparation only matters when it changes the outcome. In Firmulate’s worst-week experiment, the winning agents were not merely persuasive. They did their homework, found the fact hidden two references deep and finished the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Happy Birthday Wishes to Daughter From Father

Yearning to express your love on your daughter's special day? Dive into heartfelt birthday wishes from a father's heart.

Best Dad Good Quotes: Inspiring Words for Fathers

Discover the best dad good quotes to celebrate fatherhood. Find inspiring words to honor your dad or uplift yourself as a father. Share these heartwarming messages today

Birthday Wishes for Dad From Son: Celebrate His Day

Bring joy to your dad's birthday with heartfelt wishes that show love, gratitude, and admiration – make his day unforgettable!

Cute Father-Daughter Quotes and Sayings to Share

Peek into the tender world of father-daughter love with heartwarming quotes that will melt your heart and leave you wanting more.