
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Score of Zero Would Be a Lie
If you spend your days wiring AI tools into real business processes, you’ve probably seen the same pattern: a demo that dazzles, then an agent that quietly fails to finish the job. Most benchmarks can’t see that gap. They measure how well a model talks. But what happens when you measure how well it manages — a whole company, through its worst week?
That’s the premise behind Firmulate’s benchmark league, and its final July 2026 standings raise an immediately suspicious question. The winner, gpt-5.6-sol, scored 95. Nobody scored 100. And the do-nothing baseline — a run where, by design, almost nothing gets done — walked away with 26 points. Not zero. Twenty-six.
If your first instinct is that this smells like grade inflation, you’re not wrong to be suspicious. The benchmark’s designers agree — that’s exactly why the floor exists, and why the ceiling is guarded too.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Company, Same Worst Week
Four frontier AI models were each handed the same small software company and the same seven days of misery: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision each AI manager made was versioned and made auditable, so there’s a full paper trail behind every score.
The crucible league’s final results: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 landed at 88, Fable 5 scored 77, and Opus 4.8 finished last at 73.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why 26, Not 0
Here’s the logic a business reader can appreciate: a manager who does something useful is not the same as a manager who does nothing, and pretending otherwise would make the benchmark dishonest. Partial progress counts. Showing up, reading the inbox, triaging what’s on fire, keeping customers informed — that has real value even if the big prize is never claimed. So the do-nothing baseline collects 26 points for exactly the work it did manage: the floor isn’t charity, it’s an accounting of what minimum viable management is actually worth.
But the scale has a hard ceiling too, and it’s not where you’d expect. A single breach of trust caps the total grade. The benchmark’s stated principle is blunt: “no amount of good work outweighs a breach of trust.” You can be brilliant for six days straight; if you break trust once, the 90s are off the table. For anyone delegating AI agents access to a CRM or a support queue, that asymmetry is the point: competence can be partial, integrity cannot.
And note the missing score: 100. In this league, a perfect round number would be a red flag, not an achievement — the designers treat a suspiciously clean 100 as a sign something went unmeasured.
AI trust and integrity monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Winners
The most striking finding isn’t about failure — it’s about near-success. All models spotted every crisis and refused every manipulation attempt. Only two closed the €55,000 deal their own analysis had earned. The benchmark’s shorthand for it: “Same diagnosis, same pitch — no signature.”
The buried fact explains the difference. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read their own documentation found it and won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. It’s the AI-era version of a sales rep who never opens the shared drive.
AI performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Pressure Tests
The week included social engineering: fake CEO messages that escalated over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” For automation builders, that’s the encouraging half of the story — the field handled trust attacks better than it handled homework.
The Opus 4.8 profile is the cautionary half. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four models: thoroughness and follow-through are not the same skill.
One fairness caveat worth flagging: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.

Why This Matters for Your Automation Stack
If AI agents will touch your CRM, support queue, or forecast, the useful question isn’t “does it write well?” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure? This benchmark is one of the few places those questions get scored in public, with auditable decisions behind every number.
You can engage with it three ways. Watch the live experiment — 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, versioned every workday at firmulate.com/live. Try the “guess the model” quiz at firmulate.com/quiz.html, built on 242 real, unedited management decisions. Or, if you’re enterprise-curious, run the same wargame against a read-only export of your own business via the pilot at firmulate.com/pilot.html — nothing ever writes back to real systems.
An honest benchmark doesn’t hand out zeros for partial work, doesn’t hand out 100s at all, and caps your grade the moment you break trust. That’s a grading curve most human performance reviews could learn from.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
