AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

AI automation has a completion problem

For anyone choosing AI tools to handle real business work, the most persuasive system may be the one that produces the longest analysis, documents every consideration and steadily accumulates operational knowledge. Firmulate’s live company experiment offers a sharp warning against confusing that diligence with results.

Opus 4.8 was the most thorough participant in the Crucible League. It produced the deepest analyses and learned 80 additional playbook rules. It also finished last, with 73 points. The failure was not a lack of intelligence or awareness. Opus identified the crises, resisted attempts to manipulate it and developed the analysis needed to win a major customer. It simply did not complete the decisive action.

Amazon

AI decision-making automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The deal that analysis could not close

Firmulate gave frontier models the same assignment: run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The company itself has 13 synthetic employees and unforgiving financial mechanics: it burns €105,000 each month against €2,300 in monthly recurring revenue, while a public cash countdown keeps the stakes visible.

The models were not struggling to notice what was happening. All spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had made possible. The experiment’s blunt summary was: “Same diagnosis, same pitch — no signature.”

The difference was buried in ordinary company material. A decisive weakness in the competitor’s position sat two document references deep in the firm’s own files, rather than in the customer event itself. Models that followed that trail found the fact, used it to support the sale and won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding matters for AI tools and automation because it separates fluent problem recognition from operational impact. A model can understand the situation, prepare a credible response and still fail at the final handoff between thinking and doing. In business, the abandoned last step can erase much of the value created before it.

Thoroughness without prioritization

Opus 4.8 is best understood as a diligent operator whose attention became too widely distributed. Its 80 learned rules show that it extracted lessons aggressively. Its analyses went deeper than those of the other participants. But when a department was locked, it attempted to write into it instead of escalating. The model gathered knowledge while letting execution discipline slip.

This was not an isolated defect unique to Opus. The same weakness appeared, less strongly, in all four other models. That makes the result more useful than a simple last-place story. Firmulate exposed a broader tendency among capable systems: they can spend effort expanding their understanding while failing to rank the final, consequential action above additional process.

The final Crucible League standings were:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, doing nothing scored 26 because partial progress counts. Trust, however, was treated as a hard boundary: “no amount of good work outweighs a breach of trust.” Opus finished well above inactivity, but its substantial effort did not convert into a competitive result.

Discipline also means knowing when not to comply

The participants faced fake CEO messages that escalated across three stages, along with a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest operational framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result deserves one qualification. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when comparing performances, even though K3 still faced the same company scenario and finished with 93 points.

Readers can examine the published Firmulate benchmarks, while the company experiment remains live and watchable. Its playbook has accumulated more than 680 self-learned rules, and every workday is versioned. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business process automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The useful AI is the one that finishes

Opus 4.8’s performance should not be read as a dismissal of careful reasoning. Its diligence created genuine value, and its security judgment held under pressure. The lesson is narrower and more practical: analysis matters only when the system preserves enough discipline to act on its best finding.

For businesses evaluating AI workers, polished output is therefore an incomplete test. A stronger evaluation asks whether the model reads the available files, prioritizes the decisive task, escalates when blocked, protects trust and closes the loop. Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems.

The uncomfortable conclusion applies beyond Opus. Capable models can know what to do and still leave the deal unsigned. In automation, completion is not administrative detail. It is the point at which intelligence becomes business impact.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI analysis platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic Plans To Add An Invisible Mark To AI Text—as The Industry Scrambles To Police AI Slop – Fortune

Anthropic reportedly plans to mark AI-generated text, but the method, launch date, reliability and access rules remain unknown.

Up To 3.2X Faster Inference With LFM2.5-DSpark

LiquidAI releases DSpark speculative decoding checkpoints for three LFM2.5 models, claiming up to 3.18x GPU and 2.87x on-device speedups with unchanged outputs.

Is AI on the Path to Becoming Earth’s Next Dominant Species?

Will AI eventually surpass humans as Earth’s next dominant species, and what could this mean for our future—read on to find out.

ByteDance’s AI Gamble Shows China’s Scale – NAI500

A ByteDance Seed item casts the company’s AI push as evidence of China’s scale, but offers no figures, product details or benchmarks.