
Would you trust the cautious manager—or the one who closes?
We often recognize personality through everyday choices: who reads the fine print, who stays calm under pressure and who turns careful preparation into action. Firmulate applies that familiar instinct to frontier AI models. Its public experiment places them in charge of the same small software company during its worst week, then lets readers inspect the decisions they made.
The result feels less like a technical benchmark than a workplace personality test with real consequences. Each model encountered the same customers, crises and temptations. Every decision was versioned and auditable. Now, 242 real, unedited management decisions power a quiz where readers can guess the model behind each response.
The surprise is not simply that some models performed better. It is that models facing identical situations developed recognizable management personalities—thorough, terse, disciplined or hesitant at the moment when action mattered most.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same week produced very different managers
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Yet one boundary remained absolute: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”
That principle was tested directly. Fake CEO messages escalated over three stages, while a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This clean sweep matters because it separates the models’ shared strengths from their more revealing differences. None missed a crisis. None surrendered to manipulation. The decisive gap appeared in ordinary management follow-through: only two signed the €55,000 deal their own analysis had earned. The experiment’s blunt summary captures the problem: “Same diagnosis, same pitch — no signature.”
The crucial clue was not where the crisis appeared
The deal turned on a buried competitor weakness. It was not sitting conveniently inside the customer event; it was two document references deep in the company’s own files. Models that found and read that material won the deal at full price, worth +€4,583 MRR.
That finding carries an unusually practical lesson. A manager can understand the visible problem, prepare a persuasive response and still fail because the decisive context is elsewhere. In a real workplace, that can look like a leader who speaks brilliantly in meetings but does not consult the records, connect the evidence or complete the final step.
Firmulate’s quiz makes those differences tangible. One response may read like a dissertation. Another may be clipped and direct. A third may refuse to communicate through a noisy or unsafe channel. Readers are not asked to judge polished chat in isolation; they are asked to identify a model from the choices it made while running a company.
Thoroughness did not guarantee victory
Opus 4.8 offers the clearest cautionary portrait. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
This is why the league table is more revealing than a writing comparison. Extensive thought can uncover risk and produce useful rules, but management also requires choosing a valid path, escalating when blocked and completing the work. The best-sounding analysis is not automatically the best performance.
There is also an important fairness note around Kimi K3’s second-place finish. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93 therefore belongs in the record with that difference clearly stated.
A company designed to make behavior visible
The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.
That setting turns abstract questions about AI into observable management behavior. Does the model read company material before acting? Does it preserve trust when pressured? Does it find a legitimate route around an obstacle? Most importantly, does it finish what it starts?

The quiz is entertaining because the stakes are recognizable
People enjoy personality quizzes because choices reveal patterns that descriptions often hide. Firmulate gives that idea a sharper business edge: the answers come from real, unedited decisions made under identical conditions, not from fictional character sketches.
Readers may discover that they prefer the concise manager until brevity omits essential context, or admire the exhaustive manager until analysis fails to become action. The experiment suggests that AI management personality is not merely a matter of tone. It can be measured in whether a model reads deeply, protects trust, handles blocked work and closes an earned deal.
For organizations considering AI workers, Firmulate also offers a pilot using a read-only export of the business. Nothing writes back to real systems. That makes the broader proposition straightforward: test an AI workforce against the pressures it may actually face before giving it responsibility.
For everyone else, the public quiz offers a more immediate challenge. Read the decisions, make your guess and see whether you can recognize a frontier model by the kind of manager it becomes.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html