
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
What marketers can learn from AI’s worst week at work
For anyone managing campaigns, ecommerce operations or customer relationships, polished copy is no longer the most interesting test of artificial intelligence. The harder question is whether an AI manager can find the decisive customer insight, resist pressure to break the rules and complete the commercial task it started.
Firmulate has turned that question into an unusually revealing reader challenge. Its guess-the-model quiz presents 242 real, unedited management decisions. Readers see what an AI did in a business situation and try to identify which frontier model was responsible.
The game works because the differences are not merely stylistic. Behind every answer is a live, auditable experiment in which each model ran the same small software company through its worst week. The customers, crises and temptations remained constant. Only the model changed.
AI management decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Identical situations, different management personalities
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
Those rankings matter, but the more useful story lies inside the decisions. Every model spotted every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had justified. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction should sound familiar to commercial teams. Recognizing an opportunity, developing the right argument and finishing the sale are separate capabilities. An AI can produce impressive analysis while still failing at the moment when analysis must become accountable action.
The customer clue was not in the customer event
The most consequential information was easy to miss because it was buried in the company’s own material. A competitor weakness sat two document references deep in the files rather than appearing in the customer event. The models that followed those references won the deal at full price, adding €4,583 in monthly recurring revenue.
For marketing and ecommerce leaders, this is the experiment’s sharpest lesson. An AI’s apparent intelligence may depend on whether it reads beyond the immediate ticket, message or campaign brief. Customer context can live in sales notes, previous research, product documentation or competitive material. The decisive difference may be persistence in finding it.
The quiz makes that behavior visible without reducing it to a leaderboard. Readers can compare decisions and begin to recognize recurring tendencies: exhaustive investigation, disciplined restraint, incomplete execution or attempts to push through a blocked path instead of escalating appropriately.
Pressure revealed strong boundaries
The week also included fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is commercially significant. AI systems working around customer data, forecasts, pricing or communications will inevitably encounter requests framed as urgent, confidential or supposedly authorized by a senior executive. In this experiment, the models consistently recognized the danger and held the line.
There is an important fairness qualification when comparing the field. Kimi K3 ran without an effort parameter and therefore used the API default, while the other participants ran at xhigh. Its second-place finish should be read with that difference in mind.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against equating volume with management quality. It was the most thorough participant, learning an additional 80 playbook rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department rather than escalating. The same weakness appeared more mildly in all four of the others.
This is what gives the decisions their personality. The contrast is not simply between smart and less smart models. It is between different patterns of managerial behavior: how deeply they investigate, how reliably they respect boundaries and whether they convert sound judgment into completed work.
The surrounding company makes those choices consequential. Firmulate’s live operation has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is therefore watchable as an operating business simulation, not presented as a collection of isolated chat samples.

customer insight discovery tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The practical test is whether the work gets finished
For business leaders, the quiz is entertaining because the models are difficult to identify from individual decisions. It is useful because repeated choices expose consistent differences. The strongest participant must do more than sound competent: it must read the available material, withstand manipulation, protect trust and close the loop.
That is a better standard for evaluating AI in marketing, ecommerce and customer operations than fluency alone. Firmulate’s experiment shows that models can agree on the crisis and still diverge at the point where revenue, discipline and responsibility meet. The management personality is not merely how an answer reads. It is what the model does next.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.