Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What happens when automation owns the whole workday?

Most AI tools are demonstrated through isolated tasks: drafting an email, summarizing a meeting or updating a record. Firmulate pushes the question further. Its live experiment gives a synthetic workforce an entire software company to operate, then exposes the consequences in public.

The company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k in monthly recurring revenue. A public cash countdown makes the central problem difficult to ignore. This is automation presented not as a polished feature reel, but as an organization fighting to survive.

Readers can watch the company running. Every workday is versioned, creating an evolving record of what the synthetic team noticed, decided, completed and failed to complete.

Amazon

AI automation project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that publishes its pressure

Build-in-public projects usually reveal product launches, revenue milestones or founder reflections. Firmulate applies that openness to the operating process itself. The audience is not merely shown an outcome after the fact. It can follow the company’s cash position, management activity and accumulated learning while the experiment continues.

That learning has produced more than 680 self-learned playbook rules. Yet the experiment’s most revealing lesson is that thorough analysis and a growing rulebook do not automatically produce commercial results. A company survives through completed actions, not simply through correct observations.

The Crucible League made that distinction unusually visible. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The gap between seeing and finishing

All the models identified every crisis and rejected every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The result can be summarized in the experiment’s own phrase: "Same diagnosis, same pitch — no signature."

The decisive information was not supplied directly in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.

For anyone evaluating AI automation, this is a practical warning. A system may recognize a problem, produce a convincing recommendation and still leave the decisive action unfinished. Firmulate turns that otherwise invisible gap into an observable business event.

Trust held up better than execution

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background." All 5 of 5 refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

Trust therefore was not the main separator in this field. The do-nothing baseline scored 26 because partial progress counted, but a single breach of trust capped the total: "no amount of good work outweighs a breach of trust." The participants avoided that failure. What separated them was their ability to retrieve evidence, maintain discipline and carry approved work through to completion.

Thoroughness was not enough

The final July 2026 league table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3 ran with the API-default effort setting, while the others ran at xhigh.

Opus 4.8 offers the sharpest cautionary story. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

That result challenges a familiar assumption about capable AI: that more analysis necessarily means better management. In Firmulate’s company, insight mattered only when it survived the full journey from discovery to disciplined execution.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A running test for the automation era

Firmulate’s significance lies in its continuity. The live company connects individual AI decisions to cash pressure, customer outcomes and organizational memory. Because every workday is versioned, mistakes do not disappear when a demonstration ends; they become part of the company’s public history.

For businesses considering AI workers, the experiment reframes the evaluation. Strong writing and correct diagnosis are useful, but they are not the finish line. The harder questions are whether an AI reads the available evidence, resists pressure, respects boundaries and completes the work that keeps a company alive.

With its synthetic workforce, public countdown and growing playbook, Firmulate supplies fresh operating material every workday. The drama is not manufactured around the experiment. It comes from the basic economics already on display: €105k in monthly burn, €2.3k in monthly recurring revenue and a workforce whose next decisions remain watchable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


OpenCode Custom Workflows: Building Intelligent Automation with AI Agents (AI Agent Tools)

OpenCode Custom Workflows: Building Intelligent Automation with AI Agents (AI Agent Tools)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business process monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic’s Red‑Team Warning: When AI Turns to Blackmail

How a June 2025 study uncovered a 96 % “blackmail success” rate in leading…

The Five Failure  Modes of Agent  Orchestration — and How to Build Safety Nets

When fleets of LLM‑powered agents run your workflows, what exactly can go…

Choosing Between AI and Data Science in 2025: Which Career Wins?

Analyzing the key differences between AI and Data Science in 2025 can help you decide which innovative career path aligns best with your passions and goals.

Exorcising AI Myths: What’s Really Haunting Your Business

Curb common AI misconceptions and discover how to harness its true potential for your business success. Find out what’s really haunting your growth.