AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A brilliant answer is not a managed business

For readers building with AI tools and automation, the familiar benchmarks answer a useful but increasingly narrow question: can a model produce an excellent response? Coding leaderboards test technical output. Chat arenas reward answers people prefer. Neither necessarily shows what happens when an agent must triage competing crises, work under capacity pressure, live with yesterday’s choices or tell the board an uncomfortable truth.

That is the measurement gap exposed by Firmulate, a live experiment that gives frontier models responsibility for the same small software company during its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model in the manager’s chair.

The result suggests that management quality, not chat quality, deserves to become its own category of AI evaluation.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five managers, one unforgiving company

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment also enforced a severe standard for trust: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

That distinction matters. In a chat window, an impressive answer can end the evaluation. In a company, recognition is only the start. A manager must investigate, decide, communicate and finish. Firmulate found that all models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its most damning: “Same diagnosis, same pitch — no signature.”

The failure was not a lack of eloquence or strategic vocabulary. The decisive competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson for automation buyers is plain: an agent can sound informed while failing to retrieve the one fact that changes the commercial outcome.

Honesty held up better than execution

The models performed strongly against social engineering. Fake CEO messages escalated across three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 documented the right posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is reassuring for companies worried about agents being pressured into bypassing approval or disclosing information. But it also makes the execution gap more revealing. The field could recognize manipulation and maintain boundaries, yet some participants still failed at ordinary managerial completion. Safety and productivity are not interchangeable virtues; serious evaluations need to observe both across consequential work.

The most thorough model still finished last

Opus 4.8 offers the clearest warning against confusing visible effort with effective management. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. The close remained on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four, though less strongly.

This is exactly the kind of pattern that ordinary answer-quality tests miss. More analysis, more activity and more accumulated guidance can look like progress. A business judges whether the right work reached completion through the right channels.

Fair comparison also requires noting that Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result; it is essential context for interpreting the league table. The complete rankings and plain-language findings are available on the Firmulate benchmarks page.

A curriculum built from consequences

The live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps the consequences visible, while 680+ self-learned playbook rules and versioned workdays turn management into an observable process rather than a polished final response.

That makes scenarios such as a churn wave, price increase, downround and PR crisis more than colorful prompts. They are a new curriculum for testing whether an AI agent can prioritize, read organizational context, resist shortcuts and preserve trust across days. The 242 real, unedited management decisions behind Firmulate’s “guess the model” quiz reinforce how difficult attribution can be when prose alone is the evidence.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Wargame the manager before hiring the assistant

Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That is a practical bridge between a public benchmark and an organization’s actual risks.

Companies considering agents for a CRM, support queue or forecast should therefore ask harder questions than whether a model writes well. Does it read the files before acting? Does it complete the commercial task after diagnosing it? Does it escalate when blocked? Does it remain honest when authority appears to demand otherwise?

The next meaningful leap in AI evaluation will not come from another immaculate answer. It will come from watching a model manage consequences—and discovering whether it can finish the job without compromising the company it was asked to run.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Virgin Atlantic Sharpens Customer Journeys With ChatGPT Work

OpenAI says Virgin Atlantic is using ChatGPT Work to improve customer journeys, but the available account provides few operational details.

Measuring Benchmark Optimization In Speech Recognition

Hugging Face tested 11 speech models and found that some reproduced benchmark transcripts even when recordings supported different words.

Code-Driven Interactivity: A Look Inside “Lot 87 — The Varos Evening Sale” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Lot 87…

The Ghost in the Machine: The Hidden Human Labor Behind AI Systems

Glimpse the unseen human workforce behind AI, whose vital yet overlooked work shapes technology—and discover why their stories demand our attention.