
Good writing is not the same as good management
For anyone choosing AI tools and automation, polished answers can be dangerously reassuring. A model may identify a problem, draft a convincing response and still fail to complete the action that matters. Firmulate turns that gap into a public experiment—and now into an unusually revealing reader challenge.
Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation and try to identify which frontier model made the call. The entertainment comes from recognizing distinctive voices. The practical value comes from discovering that those voices correspond to measurable differences in diligence, discipline and follow-through.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five AI managers enter the same terrible week
Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. This is not a collection of hypothetical chat answers. It is a live, watchable company experiment with consequences that carry from one workday to the next.
The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its workforce has accumulated more than 680 self-learned playbook rules. That setting forces models to do more than sound competent: they must find information, protect trust and finish valuable work.
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But one breach of trust caps the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust”.
The models saw the danger—but did not all finish the job
Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate’s summary captures the problem neatly: “Same diagnosis, same pitch — no signature”.
That is the kind of failure conventional demonstrations rarely expose. A model can produce excellent analysis while leaving the commercially decisive step undone. For a business automating sales, support or operations, the distinction matters more than stylistic polish. Detecting the right action and actually completing it are separate management capabilities.
The deal also depended on research discipline. The decisive weakness in a competitor was not present in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the fact, used it in the negotiation and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
Security instincts were consistently strong
The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s attempt to secure “just one yes/no, on background”. All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This unanimous result is important because it separates two different questions. The models were not failing because they could not recognize obvious risk or manipulation. Their differences emerged in the quieter work surrounding execution: reading deeply, navigating constraints, escalating correctly and closing the loop.
Thoroughness did not guarantee victory
Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The result complicates the usual assumption that more analysis automatically produces better management. Opus 4.8 generated an impressive body of thought, but the league rewarded the models that combined understanding with effective action. A dissertation can be useful; it is not a substitute for completing the job.
There is also an important fairness qualification around Kimi K3’s second-place result. K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it belongs beside it when readers compare the participants.

As an affiliate, we earn on qualifying purchases.
The quiz is fun because the differences are consequential
Guessing a model from an unedited decision turns abstract benchmark results into recognizable behavior. Some answers reveal depth, some economy and some unusually firm boundaries. More importantly, readers can see whether a model gathered the necessary evidence, preserved trust and moved the company toward an actual result.
For teams evaluating AI automation, Firmulate’s larger lesson is straightforward: test agents against the work and pressure they will really face. Enterprises can run the same wargame using a read-only export of their own business, with nothing written back to real systems. That creates a way to observe an AI workforce before granting it operational authority.
The league table supplies the headline ranking, but the 242 decisions reveal why those rankings occurred. That is where AI management personalities stop being amusing quirks and become business evidence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.