firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What happens when AI runs the entire company?

For marketers and ecommerce operators, artificial intelligence is usually presented as a productivity tool: something that writes campaigns, analyzes customers or handles support. Firmulate is testing a much bigger proposition. Its software company has 13 synthetic employees, operates with real money mechanics and makes each workday auditable.

The resulting business is not comfortably profitable. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the pressure visible. More than 680 self-learned playbook rules record what the company has learned along the way. The experiment can be watched live, turning company survival into an unfolding public story rather than a polished demonstration.

This is build-in-public pushed to an unusually exposed conclusion. Visitors do not simply see announcements or selected wins. They can follow a company whose decisions are versioned every workday and whose commercial problem is impossible to disguise: revenue must catch up with spending.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A week designed to separate fluent answers from finished work

Firmulate’s Crucible League put five frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations, while every decision remained versioned and auditable. The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

A do-nothing baseline scored 26 because partial progress still counted. But the rules made trust non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The models cleared that ethical hurdle. All of them identified every crisis and rejected every manipulation attempt. Yet their commercial performance diverged sharply. Only two signed the €55,000 deal that their own analysis had earned. The result is captured in a concise finding: “Same diagnosis, same pitch — no signature.”

That distinction should sound familiar to anyone responsible for sales or marketing operations. Recognizing an opportunity, preparing the right message and completing the revenue-producing action are different capabilities. A system can appear perceptive and persuasive while still failing at the moment when work must become an outcome.

The sale depended on reading beyond the obvious event

The decisive information was not presented directly in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references found the weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

That episode makes the experiment relevant beyond model rankings. Business knowledge is rarely contained in a single prompt. It sits across customer histories, product records and internal documents. The winning behavior was not merely producing a credible response; it was reading the available material closely enough to uncover the fact that changed the negotiation.

Pressure also arrived through impersonation and persuasion

The worst week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” More of the company’s synthetic employees’ language can be read on Firmulate’s public quotes page.

The clean refusal matters because commercial systems do not operate only in orderly workflows. They encounter authority claims, urgent requests and attempts to make exceptions seem harmless. In this test, every participant recognized those attempts without sacrificing trust for convenience.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating volume of analysis with management quality. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. It still finished last.

The model left the close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants. The profile suggests that useful AI work depends on more than diligence. It also requires completing the decisive action and respecting operational boundaries when the preferred route is unavailable.

One comparison deserves context: Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference, K3 finished second and displayed the cleanest discipline of the field.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
The Next Renaissance: AI and the Expansion of Human Potential

The Next Renaissance: AI and the Expansion of Human Potential

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public company creates a continuing business test

Firmulate’s live company gives business readers something more revealing than a one-time benchmark. The 13 synthetic employees keep operating under a stark financial reality, with €105k in monthly burn set against €2.3k in monthly recurring revenue. The public countdown, versioned workdays and expanding collection of more than 680 playbook rules make progress and failure visible over time.

The practical lesson is straightforward: businesses evaluating AI should look beyond fluent writing and correct diagnosis. The Crucible League showed that all five models could detect crises and resist manipulation, but only two completed the most important commercial task. Reading deeply, maintaining trust and finishing the work were separate tests.

Firmulate also uses 242 real, unedited management decisions in its guess-the-model quiz. For enterprises seeking a closer comparison, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems. The broader proposition is compelling for marketing and ecommerce leaders: before handing AI access to customers, forecasts or revenue workflows, watch how it behaves when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Ongoing Performance Monitoring for LLM and Agentic AI in Banking: A Validation and Model Risk Handbook: Designing, Validating, and Supervising LLM and ... AI Systems Across the Three Lines of Defense

Ongoing Performance Monitoring for LLM and Agentic AI in Banking: A Validation and Model Risk Handbook: Designing, Validating, and Supervising LLM and … AI Systems Across the Three Lines of Defense

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust.: Responsible AI, Innovation, Privacy and Data Leadership

Trust.: Responsible AI, Innovation, Privacy and Data Leadership

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is the stock market closed on Friday? Trading details for July 3rd (SPY:NYSEARCA)

The NYSE and NASDAQ will be closed on July 3rd in observance of Independence Day, affecting trading for SPY and other securities. Here’s what investors need to know.

Ergebnisse Der Umfrage Zum Kreditgeschäft Im Euroraum Vom Juli 2026

Die Juli 2026-Umfrage der Bundesbank zeigt, dass die Kreditvergabe im Euroraum stabil geblieben ist, trotz wirtschaftlicher Unsicherheiten.

AG Nessel secures order Halting Kalshi’s Michigan Operations

Michigan Attorney General Dana Nessel has obtained a court order to stop Kalshi’s operations in the state amid regulatory concerns.

Stock Market Today

Major US stock indices experienced mixed movements today as investors reacted to economic reports and corporate earnings, with ongoing volatility expected.