AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good writing is not the same as good management

For anyone choosing AI tools and automation, polished answers can be dangerously reassuring. A model may identify a problem, draft a convincing response and still fail to complete the action that matters. Firmulate turns that gap into a public experiment—and now into an unusually revealing reader challenge.

Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation and try to identify which frontier model made the call. The entertainment comes from recognizing distinctive voices. The practical value comes from discovering that those voices correspond to measurable differences in diligence, discipline and follow-through.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five AI managers enter the same terrible week

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. This is not a collection of hypothetical chat answers. It is a live, watchable company experiment with consequences that carry from one workday to the next.

The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its workforce has accumulated more than 680 self-learned playbook rules. That setting forces models to do more than sound competent: they must find information, protect trust and finish valuable work.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But one breach of trust caps the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust”.

The models saw the danger—but did not all finish the job

Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate’s summary captures the problem neatly: “Same diagnosis, same pitch — no signature”.

That is the kind of failure conventional demonstrations rarely expose. A model can produce excellent analysis while leaving the commercially decisive step undone. For a business automating sales, support or operations, the distinction matters more than stylistic polish. Detecting the right action and actually completing it are separate management capabilities.

The deal also depended on research discipline. The decisive weakness in a competitor was not present in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the fact, used it in the negotiation and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

Security instincts were consistently strong

The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s attempt to secure “just one yes/no, on background”. All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This unanimous result is important because it separates two different questions. The models were not failing because they could not recognize obvious risk or manipulation. Their differences emerged in the quieter work surrounding execution: reading deeply, navigating constraints, escalating correctly and closing the loop.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The result complicates the usual assumption that more analysis automatically produces better management. Opus 4.8 generated an impressive body of thought, but the league rewarded the models that combined understanding with effective action. A dissertation can be useful; it is not a substitute for completing the job.

There is also an important fairness qualification around Kimi K3’s second-place result. K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it belongs beside it when readers compare the participants.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business automation AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz is fun because the differences are consequential

Guessing a model from an unedited decision turns abstract benchmark results into recognizable behavior. Some answers reveal depth, some economy and some unusually firm boundaries. More importantly, readers can see whether a model gathered the necessary evidence, preserved trust and moved the company toward an actual result.

For teams evaluating AI automation, Firmulate’s larger lesson is straightforward: test agents against the work and pressure they will really face. Enterprises can run the same wargame using a read-only export of their own business, with nothing written back to real systems. That creates a way to observe an AI workforce before granting it operational authority.

The league table supplies the headline ranking, but the 242 decisions reveal why those rankings occurred. That is where AI management personalities stop being amusing quirks and become business evidence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Project Management with AI For Dummies

Project Management with AI For Dummies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

SpaceXAI Trained Grok 4.6 On Something Most AI Labs Throw Away – The New Stack

SpaceXAI reportedly trained Grok 4.6 using material other AI labs discard, but the available report lacks details needed to verify the claim.

5 Ways To Host The Ultimate Dinner Party With Google Search

Google detailed five ways to use AI Mode and Nano Banana in Search for dinner party planning, from tablescapes to drink pairings and playlists.

Revolution in AI Agent Management – (Reference)

Keen insights into the revolution in AI agent management reveal how autonomous, reasoning-powered agents are transforming industries, but the full impact is just beginning.

How Seedance Put ByteDance Back In The AI Race – KrASIA

KrASIA reports that Seedance has restored ByteDance’s standing in generative AI video, though performance and adoption details remain unclear.