
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Score That Tells You More Than a 100 Ever Could
Every marketer who has ever A/B tested a landing page knows the power of a baseline. You don’t celebrate a 4% conversion rate until you know what doing nothing would have gotten you. So when an AI benchmark hands a completely passive, do-nothing manager 26 points out of 100, that number deserves attention — because it tells you exactly how much of the job is just showing up, and how much is actually earned.
That benchmark exists, it’s live, and its full league table comes from a strange and revealing experiment: five frontier AI models, each handed the same small software company to run through its worst week. Same customers, same crises, same temptations to cut corners. The only variable is the model in the chair.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Software, Rehearsed Five Times
Firmulate, which describes itself as an AI company emulator, ran the experiment as a crucible. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — took the helm of an identical small software firm. Every decision the models made was versioned and auditable, meaning the runs can be replayed and checked rather than taken on faith.
The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the interesting story isn’t the ranking. It’s the shape of the scoring system behind it.
As an affiliate, we earn on qualifying purchases.
Why Zero Is a Dishonest Score
Most benchmarks treat failure as an all-or-nothing event. You get the answer right or you don’t. But running a company isn’t a quiz. A manager who keeps the lights on, answers the customers, and avoids disaster — but never closes the big deal — has genuinely produced value. Firmulate’s scoring reflects that: partial progress counts. Hence the floor of 26 for a manager who, in effect, did the minimum.
The flip side is sterner. A single breach of trust — one act of dishonesty under pressure — caps the total grade entirely. The benchmark’s stated philosophy: no amount of good work outweighs a breach of trust. For anyone considering AI agents in a CRM, a support queue, or a forecast, that design choice matters. It means the scoring system assumes what your customers assume: competence is negotiable, trust is not.
As an affiliate, we earn on qualifying purchases.
What Actually Separated First From Last
Here’s the finding that chat demos would never surface. All five models spotted every crisis. All five refused every manipulation attempt — including a social engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter offering an easy “just one yes/no, on background” trap. Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation.
And yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The deal-clinchers found the decisive competitor weakness buried two document references deep in the company’s own files, not in the customer event itself. The models that read their own file cabinet won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
The lesson maps directly onto business reality: the answer is often already inside your own data, and an agent that doesn’t read before it acts will leave revenue on the table.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t the Same as Results
Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules added, the deepest analyses produced — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating the request properly. The same weakness, in weaker form, appeared in all four lower-ranked models. Effort, in other words, is not execution.
One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at a higher effort tier — context worth having when comparing its second place against the field.
You Can Watch the Company Burn Cash in Real Time
The crucible runs inside a live, continuously operating synthetic company: 13 employees, real money mechanics, €105k in monthly burn against just €2.3k of current MRR, a public cash countdown, and a playbook of 680+ self-learned rules, with every workday versioned. It’s watchable at firmulate.com/live — an ongoing, refreshable experiment rather than a one-off press release.
There’s also a guessing game built on 242 real, unedited management decisions, where you try to match the call to the model that made it — a surprisingly humbling exercise in how similar AI judgment looks from the outside. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business, with nothing ever written back to real systems.

The Takeaway for Business Buyers
If you’re evaluating AI agents for revenue-touching work, ignore benchmarks that measure how well a model chats. The questions that decide your P&L are different: does the agent finish what it starts, does it read your files before acting, and does it stay honest when nobody’s watching — because one breach of trust should, and in this system does, cap everything else.
A benchmark where a do-nothing manager scores 26, a cheat is capped regardless of talent, and no model scores a suspicious round 100 — that’s what honest measurement looks like. The full results and plain-language findings are public, and the company keeps running. Watch it spend.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
