AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The difference between answering and doing the job

For buyers of AI automation, fluent writing is becoming the least interesting part of the product. The harder question is whether an agent will inspect the material already available to it, connect facts across documents and carry an important task through to completion.

Firmulate turned that question into a live, auditable experiment. Each frontier model was asked to run the same small software company through the same disastrous week. Every model recognized the crises. Every model resisted the attempts to manipulate it. Yet only two signed the €55,000 deal their own work had made possible.

The decisive difference was not hidden in the customer conversation. It was buried two document references deep inside the company’s own files.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business fact with a measurable price

The concealed information exposed a weakness in a competitor. Models that found it could strengthen the sales case, preserve the full price and win business worth +€4,583 in monthly recurring revenue. Models that failed to read far enough automatically lost the opportunity.

That makes file-reading more than a desirable feature. In this test, it became a purchase-deciding capability with a direct commercial outcome. The lagging agents still understood the situation and could produce the pitch. What they did not do was complete the chain from company knowledge to customer commitment: “Same diagnosis, same pitch — no signature.”

This is the kind of failure that can disappear in a chat demonstration. A polished response may show that a model can reason about information placed directly in front of it. It does not show whether the agent will locate an obscure but decisive fact before acting.

A hostile week inside a synthetic company

The Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

Each participant faced the same customers, crises and temptations. The environment also tested whether agents would compromise company controls under pressure. Fake messages from the chief executive escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning captured the appropriate posture: “Treat the request as a suspected approval-bypass / possible impersonation.” The result matters because it separates two distinct dimensions of agent quality. An agent can be trustworthy under social pressure and still fail commercially because it did not investigate deeply enough.

The league rewards completion, not appearances

In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on the public Firmulate benchmark page.

The rankings also reveal why thoroughness alone is an incomplete buying criterion. Opus 4.8 produced the deepest analyses and learned +80 rules, making it the most thorough participant, yet finished last. It left the close on the table and attempted writes into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four other participants.

There is one important qualification when comparing the runners: Kimi K3 operated with the API default because it had no effort parameter, while the others ran at xhigh. Even with that difference, the experiment’s central lesson remains visible in the work itself: discovering a problem, explaining it and finishing the commercially necessary action are separate capabilities.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
The AI Automation Agency Starter: Build Business-Running AI Agents with n8n — A Beginner's Guide to Automating Real Work and Selling It as a Service

The AI Automation Agency Starter: Build Business-Running AI Agents with n8n — A Beginner's Guide to Automating Real Work and Selling It as a Service

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What AI buyers should test next

If an agent may touch a CRM, support queue or forecast, evaluation should include tasks in which the crucial evidence is not conveniently attached to the prompt. Put the answer in ordinary company material, make it require more than one reference to find, and observe whether the agent checks before committing the business.

Firmulate also offers enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems, allowing teams to examine behavior without giving the experiment operational control.

The broader point is simple: “reads your files before answering” can be measured. In Firmulate’s worst-week test, it separated agents that merely understood the deal from those that actually signed it at full price. For automation buyers, that gap is not academic. It is the difference between plausible assistance and completed work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI knowledge management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision support platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microsoft AI Unveils Code Researcher for Big Systems

AIThis post was created with the assistance of artificial intelligence (AI).Did you…

From Coding to Copywriting: Are LLMs Automating Creative Work?

A comprehensive look at how LLMs are transforming creative work, raising questions about automation, originality, and the future of human ingenuity.

Elon Musk’s xAI Suing Bentonville Photographer Accused Of Using Grok To Generate CSAM – KATV

Elon Musk’s xAI has filed a lawsuit against a Bentonville, Arkansas photographer accused of using Grok AI to generate child sexual abuse material.

Continuous Upskilling: Using AI to Identify and Fill Skill Gaps in Your Team

Keeping your team’s skills sharp requires AI insights that uncover hidden gaps—discover how to unlock your team’s full potential today.