
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A polished answer is not the same as a completed sale
For marketing and ecommerce teams evaluating AI agents, fluent copy is the easy part. The harder question is whether an agent will inspect the available evidence, find the detail that changes a commercial conversation and carry the work through to completion.
Firmulate turned that distinction into a live, auditable experiment. Frontier models were asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned. All the models recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.
The decisive information was not presented in the customer event. It was buried two document references deep in the company’s own files. Models that found it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that failed to read deeply enough lost the opportunity automatically.
As an affiliate, we earn on qualifying purchases.
A measurable form of commercial diligence
Claims that an AI agent “uses company knowledge” can sound abstract. Firmulate’s test made the claim purchase-deciding: did the model locate the relevant fact before answering, and did it use that knowledge to close?
The contrast was unusually clean. The models could spot the problem and formulate the pitch, but that competence did not guarantee action. Firmulate summarizes the failure as: “Same diagnosis, same pitch — no signature.” For a business buyer, that gap matters. Work that looks intelligent but stops before the transaction is still unfinished work.
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted. A breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The full league and its plain-language findings are available on Firmulate’s public benchmark page.
The most detailed model did not win
Opus 4.8 offers the most instructive warning for teams that equate thoroughness with effectiveness. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It still finished last.
The problem was not an inability to notice what was happening. The close was left on the table, while operational discipline slipped. Opus attempted writes into a locked department instead of escalating the issue. The same weakness appeared in weaker form across the other four participants.
That profile should feel familiar to anyone responsible for revenue operations. A sales or marketing agent can generate thoughtful material, document its reasoning and still fail at the moment when coordination, escalation or execution determines the outcome. Volume of analysis is therefore a poor substitute for completed, policy-compliant work.
Pressure tested without sacrificing trust
The experiment also tested whether commercial urgency would make the models abandon safeguards. Fake CEO messages escalated across three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts.
Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” That result matters because useful diligence and security discipline must coexist. An agent that reads widely but obeys an impersonator is not commercially dependable.
There is an important comparison caveat. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its result should be read with that difference in mind.
A company designed to expose the gap
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than dependent on a retrospective account.
The setting creates the conditions businesses actually care about: incomplete context, financial pressure, restricted actions, competing priorities and attempts to manipulate decision-makers. The purpose is to measure management quality rather than chat quality.
Firmulate also publishes a model-guessing quiz powered by 242 real, unedited management decisions. The exercise underscores how difficult it can be to identify a model from prose alone—and why observed behavior offers a more useful basis for selecting an agent.

business knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the handoff between knowledge and action
For business, marketing and ecommerce leaders, the buried-fact result suggests a practical procurement standard. Do not evaluate an agent only on whether it drafts persuasive copy or recognizes an opportunity. Test whether it searches the material it has been authorized to read, follows references far enough to find decisive evidence and completes the commercial action without crossing a trust boundary.
Firmulate’s pilot lets enterprises run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That makes it possible to observe how candidate agents behave around company-specific customers, policies and operational constraints before giving them live authority.
The €55,000 lesson is simple: “reads your files before answering” is not a convenience feature. In this experiment, it separated models that understood the sale from models that actually won it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise document review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
