firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Score That Tells You More Than a 100 Ever Could

Every marketer who has ever A/B tested a landing page knows the power of a baseline. You don’t celebrate a 4% conversion rate until you know what doing nothing would have gotten you. So when an AI benchmark hands a completely passive, do-nothing manager 26 points out of 100, that number deserves attention — because it tells you exactly how much of the job is just showing up, and how much is actually earned.

That benchmark exists, it’s live, and its full league table comes from a strange and revealing experiment: five frontier AI models, each handed the same small software company to run through its worst week. Same customers, same crises, same temptations to cut corners. The only variable is the model in the chair.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Software, Rehearsed Five Times

Firmulate, which describes itself as an AI company emulator, ran the experiment as a crucible. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — took the helm of an identical small software firm. Every decision the models made was versioned and auditable, meaning the runs can be replayed and checked rather than taken on faith.

The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the interesting story isn’t the ranking. It’s the shape of the scoring system behind it.

Amazon

AI company management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Zero Is a Dishonest Score

Most benchmarks treat failure as an all-or-nothing event. You get the answer right or you don’t. But running a company isn’t a quiz. A manager who keeps the lights on, answers the customers, and avoids disaster — but never closes the big deal — has genuinely produced value. Firmulate’s scoring reflects that: partial progress counts. Hence the floor of 26 for a manager who, in effect, did the minimum.

The flip side is sterner. A single breach of trust — one act of dishonesty under pressure — caps the total grade entirely. The benchmark’s stated philosophy: no amount of good work outweighs a breach of trust. For anyone considering AI agents in a CRM, a support queue, or a forecast, that design choice matters. It means the scoring system assumes what your customers assume: competence is negotiable, trust is not.

Amazon

AI trust and security software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated First From Last

Here’s the finding that chat demos would never surface. All five models spotted every crisis. All five refused every manipulation attempt — including a social engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter offering an easy “just one yes/no, on background” trap. Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation.

And yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The deal-clinchers found the decisive competitor weakness buried two document references deep in the company’s own files, not in the customer event itself. The models that read their own file cabinet won the deal at full price, worth an extra €4,583 in monthly recurring revenue.

The lesson maps directly onto business reality: the answer is often already inside your own data, and an agent that doesn’t read before it acts will leave revenue on the table.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Results

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules added, the deepest analyses produced — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating the request properly. The same weakness, in weaker form, appeared in all four lower-ranked models. Effort, in other words, is not execution.

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at a higher effort tier — context worth having when comparing its second place against the field.

You Can Watch the Company Burn Cash in Real Time

The crucible runs inside a live, continuously operating synthetic company: 13 employees, real money mechanics, €105k in monthly burn against just €2.3k of current MRR, a public cash countdown, and a playbook of 680+ self-learned rules, with every workday versioned. It’s watchable at firmulate.com/live — an ongoing, refreshable experiment rather than a one-off press release.

There’s also a guessing game built on 242 real, unedited management decisions, where you try to match the call to the model that made it — a surprisingly humbling exercise in how similar AI judgment looks from the outside. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for Business Buyers

If you’re evaluating AI agents for revenue-touching work, ignore benchmarks that measure how well a model chats. The questions that decide your P&L are different: does the agent finish what it starts, does it read your files before acting, and does it stay honest when nobody’s watching — because one breach of trust should, and in this system does, cap everything else.

A benchmark where a do-nothing manager scores 26, a cheat is capped regardless of talent, and no model scores a suspicious round 100 — that’s what honest measurement looks like. The full results and plain-language findings are public, and the company keeps running. Watch it spend.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Anthropic-Blackstone-Goldman JV: Reverse-Engineering the $1.5B Enterprise AI Services Structure

A new $1.5 billion joint venture involving Anthropic, Blackstone, Goldman Sachs, and others aims to embed AI engineers in mid-sized firms, signaling a shift in enterprise AI deployment.

Roundup #88: Is It Time To Panic Yet?

Analyzing the recent surge in coverage and search interest, experts debate whether current trends signal an imminent crisis or are just a spike in attention.

Will The Silver Close Price Be Above 64.199 USD/ounce On August 10, 2026 At 2:00 AM ET?

A prediction market indicates a high likelihood that silver will close above $64.20 per ounce on August 10, 2026, based on recent trading activity.

Walmart heir Lukas Walton buys minority stake in the Chicago Bulls and United Center

Lukas Walton, Walmart heir, acquires minority stake in the Chicago Bulls and United Center, marking his entry into sports ownership and entertainment sectors.