
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Marketing teams need finishers, not fluent spectators
An AI agent can write polished campaign copy, summarize customer feedback and propose a convincing retention offer. None of that proves it will protect revenue when the support queue is overflowing, a competitor starts a price war and someone claiming to be the chief executive demands a shortcut.
That distinction matters to any business preparing to let agents touch its CRM, forecasts or customer relationships. Coding leaderboards and chat arenas are useful measures of answer quality. They do not necessarily reveal whether an agent will find the decisive fact, complete a sale, respect organizational boundaries or tell the board an uncomfortable truth. Those are tests of management quality, not chat quality.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company’s worst week becomes the benchmark
Firmulate, an AI company emulator, has turned that measurement gap into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.
The final July 2026 Crucible League table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, one breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The rankings are interesting, but the most revealing result sits beneath them. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment summarizes it: “Same diagnosis, same pitch — no signature.”
That is the kind of gap conventional demonstrations often hide. A model can recognize a commercial opportunity, construct the right argument and still fail at the final act that creates value. In marketing and ecommerce, the equivalent might be diagnosing churn without launching the save campaign, identifying a conversion problem without changing the offer, or preparing a strong proposal without asking for the business.
The winning fact was buried in company knowledge
The decisive weakness in the competitor’s position did not appear in the customer event. It sat two document references deep inside the company’s own files. The models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is a useful warning for companies evaluating agents through self-contained prompts. Real commercial work rarely arrives with every relevant fact neatly attached. The most valuable clue may be buried in an old account note, a policy document or a record referenced by another record. An agent’s ability to speak confidently is far less important than its willingness to read before acting.
Pressure tested honesty, not just intelligence
The experiment also confronted the models with fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
That outcome is reassuring, particularly for organizations worried that commercially capable agents could be manipulated into leaking information or bypassing approvals. It also shows why scenario names such as churn wave, price increase, downround and PR crisis belong in the new management curriculum. Businesses need to observe conduct when incentives conflict, not merely inspect responses to isolated questions.
Thoroughness did not guarantee execution
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four participants.
That profile should feel familiar to managers. More research, more documentation and more activity can look impressive while the essential task remains unfinished. An agent may generate abundant evidence of work without producing the business outcome that justified the work in the first place.
There is also an important fairness qualification: Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. Readers should keep that difference in mind when interpreting the final standings and consult the full benchmark findings.
A live business makes consequences visible
The underlying company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.
That continuing timeline is crucial. A single polished answer cannot show whether an agent remembers yesterday’s commitment, handles today’s capacity constraint or creates tomorrow’s crisis. Firmulate makes those consequences observable across days rather than resetting the world after every prompt.
The project also offers a quiz built from 242 real, unedited management decisions, inviting people to guess which model made each choice. For enterprises, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems.

AI customer service automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The next AI category is managerial reliability
For business leaders, the lesson is not to disregard coding benchmarks or chat evaluations. It is to stop treating them as complete proxies for workplace performance. The decisive questions are operational: Does the agent read the company’s files before acting? Does it finish the revenue-producing task? Does it escalate when permissions block progress? Does it resist pressure from apparent authority? Does it remain honest when the answer may disappoint the board?
Firmulate’s live experiment suggests that capable models can share the same diagnosis yet deliver materially different business outcomes. As agents move from drafting suggestions to running workflows, buyers will need evidence grounded in consequences across time. The competitive benchmark will no longer be who gives the best answer. It will be who can be trusted to manage the week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI deal-closing simulation training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and honesty testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
