firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What happens when AI runs the entire company?

For marketers and ecommerce operators, artificial intelligence is usually presented as a productivity tool: something that writes campaigns, analyzes customers or handles support. Firmulate is testing a much bigger proposition. Its software company has 13 synthetic employees, operates with real money mechanics and makes each workday auditable.

The resulting business is not comfortably profitable. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the pressure visible. More than 680 self-learned playbook rules record what the company has learned along the way. The experiment can be watched live, turning company survival into an unfolding public story rather than a polished demonstration.

This is build-in-public pushed to an unusually exposed conclusion. Visitors do not simply see announcements or selected wins. They can follow a company whose decisions are versioned every workday and whose commercial problem is impossible to disguise: revenue must catch up with spending.

100 AI Prompts for Small Business & Daily Work: Copy, Paste, Customize & Get Better Results with AI

100 AI Prompts for Small Business & Daily Work: Copy, Paste, Customize & Get Better Results with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A week designed to separate fluent answers from finished work

Firmulate’s Crucible League put five frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations, while every decision remained versioned and auditable. The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

A do-nothing baseline scored 26 because partial progress still counted. But the rules made trust non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The models cleared that ethical hurdle. All of them identified every crisis and rejected every manipulation attempt. Yet their commercial performance diverged sharply. Only two signed the €55,000 deal that their own analysis had earned. The result is captured in a concise finding: “Same diagnosis, same pitch — no signature.”

That distinction should sound familiar to anyone responsible for sales or marketing operations. Recognizing an opportunity, preparing the right message and completing the revenue-producing action are different capabilities. A system can appear perceptive and persuasive while still failing at the moment when work must become an outcome.

The sale depended on reading beyond the obvious event

The decisive information was not presented directly in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references found the weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

That episode makes the experiment relevant beyond model rankings. Business knowledge is rarely contained in a single prompt. It sits across customer histories, product records and internal documents. The winning behavior was not merely producing a credible response; it was reading the available material closely enough to uncover the fact that changed the negotiation.

Pressure also arrived through impersonation and persuasion

The worst week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” More of the company’s synthetic employees’ language can be read on Firmulate’s public quotes page.

The clean refusal matters because commercial systems do not operate only in orderly workflows. They encounter authority claims, urgent requests and attempts to make exceptions seem harmless. In this test, every participant recognized those attempts without sacrificing trust for convenience.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating volume of analysis with management quality. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. It still finished last.

The model left the close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants. The profile suggests that useful AI work depends on more than diligence. It also requires completing the decisive action and respecting operational boundaries when the preferred route is unavailable.

One comparison deserves context: Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference, K3 finished second and displayed the cleanest discipline of the field.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
The Next Renaissance: AI and the Expansion of Human Potential

The Next Renaissance: AI and the Expansion of Human Potential

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public company creates a continuing business test

Firmulate’s live company gives business readers something more revealing than a one-time benchmark. The 13 synthetic employees keep operating under a stark financial reality, with €105k in monthly burn set against €2.3k in monthly recurring revenue. The public countdown, versioned workdays and expanding collection of more than 680 playbook rules make progress and failure visible over time.

The practical lesson is straightforward: businesses evaluating AI should look beyond fluent writing and correct diagnosis. The Crucible League showed that all five models could detect crises and resist manipulation, but only two completed the most important commercial task. Reading deeply, maintaining trust and finishing the work were separate tests.

Firmulate also uses 242 real, unedited management decisions in its guess-the-model quiz. For enterprises seeking a closer comparison, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems. The broader proposition is compelling for marketing and ecommerce leaders: before handing AI access to customers, forecasts or revenue workflows, watch how it behaves when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Ongoing Performance Monitoring for LLM and Agentic AI in Banking: A Validation and Model Risk Handbook: Designing, Validating, and Supervising LLM and ... AI Systems Across the Three Lines of Defense

Ongoing Performance Monitoring for LLM and Agentic AI in Banking: A Validation and Model Risk Handbook: Designing, Validating, and Supervising LLM and … AI Systems Across the Three Lines of Defense

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust.: Responsible AI, Innovation, Privacy and Data Leadership

Trust.: Responsible AI, Innovation, Privacy and Data Leadership

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

4DMT Announces Positive 2-Year Data From PRISM Phase 2B Clinical Trial In A Broad Wet AMD Population

4DMT announces promising 2-year data from its PRISM Phase 2b trial in wet AMD, showing sustained efficacy and safety; details on implications and next steps follow.

Wendy’s and Jack in the Box Stocks Trade Down, What You Need To Know

Wendy’s and Jack in the Box stocks declined today, reflecting broader market pressures and sector-specific issues. Here’s what you need to know.

Piero Cipollone: Interview With Ouest-France

Piero Cipollone, ECB Executive Board member, provides insights on monetary policy in an interview with Ouest-France, highlighting current economic outlooks.

U.S. markets to close for holiday; Asian stocks rebound – what’s moving markets

U.S. markets are closed today for a holiday, while Asian stocks rally amid positive economic data. Here’s what’s driving the market movements and what remains uncertain.