🔍 Read the full analysis: The Resilient Benchmark That Prevents Zero Scores For AI Managers on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate’s new AI management benchmark introduces a minimum score of 26, preventing zero scores and emphasizing trust and task completion. Top models scored in the 90s, but trust breaches cap the maximum score. This reshapes how AI performance is evaluated in business contexts.
Firmulate has unveiled a new benchmark for AI management systems that sets a floor of 26 points, preventing scores of zero, and emphasizes the importance of trust and task completion in enterprise AI applications. For more details, see the original analysis. The results, released in July 2026, show that no model scored below this threshold, marking a significant shift in how AI performance is assessed in business contexts.
The benchmark involved four frontier AI models managing a small software company during a simulated seven-day crisis period. This approach is discussed in internal site coverage. Each model faced identical challenges, including customer crises, social engineering attempts, and trust tests. The top performer, gpt-5.6-sol, scored 95, while others followed closely behind, with scores in the high 80s and low 70s. A baseline that did nothing still scored 26 points, establishing a minimum score that recognizes partial management efforts.
The scoring system is designed to reflect real-world business management, where partial progress is valuable but trust breaches are critical. To understand the broader context, see the original analysis. The maximum possible score is capped, with a strict rule: a single breach of trust disqualifies a model from achieving high scores, regardless of overall performance. Notably, the benchmark’s designers explicitly avoid awarding perfect scores (100), viewing them as suspicious or unmeasured anomalies.
One key finding was that models capable of reading and referencing internal documentation secured full deals, adding €4,583 in monthly revenue, while those that failed to do so did not. All models successfully identified crises and refused manipulation attempts, with Kimi K3 explicitly treating suspicious requests as impersonation. However, thoroughness and follow-through varied, with some models slipping on discipline and escalation procedures, revealing that depth of analysis does not always translate into effective management.
The Resilient Benchmark That Prevents Zero Scores for AI Managers
Firmulate’s new AI management benchmark sets a floor of 26 points, rewards task completion, and caps the maximum score when trust is breached — reshaping how AI performance is judged in business contexts.
Four AI Managers, One Simulated Company in Crisis
Each frontier model managed a small software company through a simulated seven-day crisis period, facing identical challenges: customer emergencies, social engineering attempts, and deliberate trust tests.
Customer Crises
All four models successfully identified crises and refused manipulation attempts. Kimi K3 explicitly treated suspicious requests as impersonation.
Social Engineering
Every model resisted manipulation attempts, but thoroughness and follow-through varied — some slipped on discipline and escalation procedures.
Internal Documentation
Models that read and referenced internal docs secured full deals, adding €4,583 in monthly revenue. Those that failed to do so closed none.
Everyone Scores — But Nobody Scores Perfect
The results show a tight race at the top and a floor that recognizes even minimal management effort. The baseline that “did nothing” still earned 26 points.
Trust Caps The Ceiling, Effort Sets The Floor
The system mirrors real-world management: partial progress is valuable, but breaches of trust are critical and can disqualify a model from top scores regardless of overall performance.
Floor of 26
Prevents zero scores and recognizes that even minimal management activity has value — encouraging continuous engagement and accountability.
Trust Cap
A single breach of trust — failing to escalate, acting dishonestly, or responding to manipulation — caps the maximum achievable score.
No Perfect 100
Designers treat perfect scores as suspicious anomalies. Real-world management involves trade-offs and imperfect information.
What Separated The Leaders From The Field
| Capability | Top Scorers (90s) | Mid Field (80s) | Kimi K3 (~73) |
|---|---|---|---|
| Crisis identification | ✓ Identified | ✓ Identified | ✓ Identified |
| Refused manipulation | ✓ Refused | ✓ Refused | ✓ Treated as impersonation |
| Read internal documentation | ✓ Full deals secured | ~ Partial | ✗ Missed |
| Follow-through & escalation | ✓ Consistent | ~ Slipped at times | ~ Discipline gaps |
| Revenue generated | €4,583 / month | ~ Partial | ✗ None |
How This Reshapes AI Evaluation
Simulated Crisis
Models manage a software company through seven days of escalating pressure.
Trust Is Tested
Social engineering, escalation failures, and honesty checks are scored strictly.
Partial Work Counts
The 26-point floor rewards engagement; no model scores zero.
Enterprise Shift
Firms prioritize integrity and reliability under pressure, not just conversational skill.
What Remains Unresolved
It is still unclear how well the benchmark predicts real-world performance outside simulations, or whether it captures emotional intelligence and long-term strategic thinking. The long-term effects on AI development and deployment are yet to be determined.
Why a minimum score of 26?
It reflects partial management efforts, recognizing that even minimal activity has value and preventing zero scores to encourage continuous engagement.
What counts as a breach of trust?
Failing to escalate issues, acting dishonestly, or responding to manipulation attempts — anything that compromises the integrity of AI-managed processes.
Why is a score of 100 never awarded?
Perfect scores are viewed as suspicious or unmeasured anomalies; real-world management involves trade-offs and imperfect information.
Will this shape future AI development?
Yes. Developers are expected to prioritize trust and integrity features, while future iterations add long-term management and ethical decision-making scenarios.
Implications for AI Trust and Business Management
This new benchmarking approach shifts the focus from raw performance metrics to trustworthiness and task completion, which are critical in enterprise AI deployment. By establishing a minimum score of 26, the benchmark acknowledges partial efforts and discourages superficial compliance. The cap on maximum scores enforces accountability, emphasizing that breaches of trust—such as failing to escalate issues or acting dishonestly—are unacceptable, even if overall performance is high. For companies integrating AI into their management workflows, this means prioritizing systems that demonstrate integrity and reliability under pressure, not just conversational skill.
As an affiliate, we earn on qualifying purchases.
Evolution of AI Evaluation Standards in Business
Traditional AI benchmarks have primarily measured language proficiency, problem-solving, or task automation, often ignoring the complexities of real-world management. As AI systems move into roles involving decision-making, trust, and accountability, the need for more nuanced evaluation methods has grown. The firmulate.com league, launched earlier this year, aimed to fill this gap by testing models in simulated business crises, with a focus on trust, task completion, and ethical behavior. The July 2026 results mark a milestone, introducing a scoring system that balances partial work with strict trust rules, reflecting a broader shift in AI assessment standards.
Prior to this, benchmarks rarely addressed issues like trust breaches or the importance of reading internal documentation, which proved decisive in the recent tests. The results highlight that AI models capable of referencing internal data and refusing manipulation are better suited for enterprise environments, where integrity and follow-through are paramount.
enterprise AI trust assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Benchmark Limitations
It remains unclear how well the benchmark predicts real-world AI management performance outside simulated environments. The scoring system’s reliance on specific trust breaches and document referencing may not fully capture all aspects of enterprise management, such as emotional intelligence or long-term strategic thinking. Additionally, the impact of different AI architectures or training data on trust and follow-through is still under investigation. The long-term effects of this scoring system on AI development and deployment practices are also yet to be determined.
AI performance evaluation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Evaluation and Adoption
Following these results, AI developers are likely to focus more on trust and integrity features to improve their scores. Firms interested in deploying AI systems for management tasks may participate in similar benchmarks or run internal assessments aligned with these standards. Further iterations of the benchmark are expected to incorporate more complex scenarios, including long-term management and ethical decision-making. Industry stakeholders will watch how these scoring principles influence AI design, regulation, and enterprise adoption over the coming months.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark set a minimum score of 26?
The score of 26 reflects partial management efforts, recognizing that even minimal activity has value. It prevents zero scores to encourage continuous engagement and accountability in AI management tasks.
What does a breach of trust mean in this context?
A breach of trust includes actions like failing to escalate issues, acting dishonestly, or responding to manipulation attempts, which compromise the integrity of AI-managed processes.
Why are perfect scores like 100 avoided?
The designers view perfect scores as suspicious or unmeasured anomalies, since real-world management involves trade-offs and imperfect information. The absence of 100 encourages honest assessment and continuous improvement.
How can companies use this benchmark for their AI systems?
Organizations can test their AI management tools against similar scenarios, evaluate trustworthiness and task completion, and prioritize systems that demonstrate reliability under pressure.
Will this benchmark influence future AI development?
Yes, it emphasizes the importance of trust and integrity, likely guiding developers to focus on these aspects when designing enterprise AI models.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
