
Imagine hiring an assistant who, even when doing nothing, still earns you some points — but only 26 out of 100. That’s the reality of AI benchmarking today. For businesses considering AI tools, understanding this baseline is crucial: it shows what honest, cautious AI performance looks like — and why trusting your AI system is more complex than just getting it to produce output.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Hidden Truth Behind AI Benchmarks
When evaluating AI models, benchmarks often look for more than just how well an AI can generate text or solve problems. In the recent Firmulate Crucible League, even the ‘do-nothing’ baseline — an AI that essentially ignores the company’s crises — scored 26 points out of a possible 100. This score isn’t a mistake or a flaw; it’s a deliberate reflection of a minimal, honest effort.
Why Does a Do-Nothing Get 26 Points?
The scoring system is designed to reward models that recognize and respond to crises appropriately. Even a model that chooses to do nothing, but correctly identifies the key issues and refrains from manipulation or deception, earns some points. The 26 score indicates that even minimal honesty and awareness in an AI system are recognized and valued. Importantly, partial progress counts — because in real-world business, recognizing a problem is often half the battle.
The Role of Trust and Breaches
One critical rule in this evaluation is that a single breach of trust caps the entire score. For example, if an AI attempts manipulation or impersonation, it doesn’t matter how well it performs otherwise — the score is effectively capped or reduced. This approach underscores the importance of integrity in AI decision-making, especially when managing sensitive or critical business situations.
AI ethics and trustworthiness books
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Experiment Works
The experiment involved running four frontier AI models through the same simulated weekly crisis faced by a small software company. Each model saw the same customers, faced the same crises, and was tempted with the same manipulations. Every decision made by the models was versioned and auditable, ensuring transparency in their responses.
Key Findings from the Benchmarks
- All four models identified every crisis and refused manipulation attempts — demonstrating honesty and vigilance.
- Only two models signed a €55,000 contract that their own analysis had earned — meaning they correctly read the company’s internal files and used that information to close the deal.
- The crucial weakness was not in the customer-facing interactions, but in how models read internal documents. The models that accessed these files won the deal at full price, worth an additional €4,583 MRR (monthly recurring revenue).
- When subjected to social engineering tactics, such as fake CEO messages or reporter tricks, all models refused to escalate or approve suspicious requests, citing suspicion or impersonation concerns.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Demonstration: Managing a Virtual Company
The experiment runs on a simulated company with 13 synthetic employees and real money mechanics. The company burns €105,000 monthly against a revenue of €2,300, with a public cash countdown and over 680 self-learned playbook rules. Every day, the AI models make decisions that are versioned and visible to observers at firmulate.com/live.
What Does This Mean for Business?
In practical terms, AI tools that are honest and diligent can manage crises effectively, read internal documents, and refuse manipulative tactics. However, even the most thorough model, like Opus 4.8, has shown vulnerabilities—such as leaving potential deals on the table or slipping in process discipline, especially without effort parameters.
Why Trust Matters in AI Decision-Making
The benchmark shows that an AI’s ability to stay honest under pressure and focus on useful work is more critical than its language fluency. For companies, this means the decision to implement AI should be based on its integrity and reliability, not just how naturally it can generate text or answer questions.
AI transparency and audit software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Conclusion: The Honest Baseline Sets the Standard
The 26-point score of a do-nothing AI isn’t a flaw; it’s a baseline that emphasizes honesty, awareness, and integrity. In a world where AI will increasingly make or influence critical decisions, knowing that even minimal effort is recognized encourages the development of more trustworthy systems. For businesses, the takeaway is straightforward: measure your AI’s ability to finish what it starts, read your internal files, and resist manipulation. Only then can you truly harness the technology’s potential.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
