
When work gets stressful, a colleague’s best idea is only useful if they follow through—and know where the boundaries are. That’s true whether the pressure comes from a difficult customer, a tempting shortcut, or a message that seems to come from the boss.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts AI models into that kind of pressure cooker. Its live experiment follows a synthetic company through real money mechanics and escalating decisions, turning the question of whether an AI can help into a more practical one: how does it behave when the week goes wrong?
A shared week, different outcomes
In the final Crucible League in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises, and temptations. The published order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s rule is plain: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
The headline finding was less about spotting trouble than acting on what the models already knew. Every model identified every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. For a business owner, that gap between sound advice and a completed decision may matter more than a polished answer in a chat window.
Top picks for "colleague trust week"
As an affiliate, we earn on qualifying purchases.
The detail hidden in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a small but telling story: useful business judgment can depend on finding the relevant detail in the material a company already has, then carrying the decision through.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In the final league, caution was visible across the field; so was the challenge of pairing that caution with effective follow-through.
Thoroughness isn’t the whole job
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. The lesson is not that analysis lacks value. It is that a company also needs an AI workforce to recognize when it should act, when it should stop, and when it should ask a person to take over.
There is a fairness detail alongside the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are also grounded in an experiment people can watch. Firmulate’s live company has 13 synthetic employees, a public cash countdown, and 680+ self-learned playbook rules. It burns €105k/month against €2.3k MRR, and every workday is versioned. The site also offers a “guess the model” quiz based on 242 real, unedited management decisions. Visit Firmulate to follow the experiment or try the quiz.
From watching to your own pilot
For an enterprise, the next step can be a wargame built around its own business. Firmulate’s pilot uses a read-only export to create a digital twin, then tests crisis scenarios against that company and produces a board report with a model ranking and the weak points in its playbooks. The stated boundary is important: nothing writes back to real systems. The exercise lets leaders examine how AI models handle their customers, pipeline, rules, and pressure before deciding how those models should be used in day-to-day work.

A model that sees the crisis but leaves the earned deal unsigned has shown both promise and a gap worth understanding. Firmulate makes that behavior visible in a live experiment—and offers enterprises a way to explore it against their own business. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
