
In the world of AI, reading more than just the surface can make or break a deal. Imagine an AI that doesn’t just respond but deeply understands your files — even two references deep — before making a decision. This ability to ‘read between the lines’ is proving to be a game-changer, especially when stakes are high.
The Experiment: Testing AI Under Real-World Pressure
Recently, four frontier AI models faced off in a unique challenge: managing a small software company’s worst week. They navigated the same crises, dealt with the same customers, and faced the same temptations to cut corners. Each model’s decisions were fully documented and auditable, ensuring transparency in performance. This rigorous setup aimed to uncover a crucial question: how well do AI agents genuinely understand and follow complex instructions?
As an affiliate, we earn on qualifying purchases.
The Key Finding: Deep File Reading Wins Deals
Results showed that all four models successfully identified every crisis and refused manipulative challenges. However, only two managed to finish the task and secure a €55,000 deal — the amount representing real revenue for the simulated company. Interestingly, the decisive advantage was linked to how deeply each model could understand the company’s internal documents.
enterprise AI file comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Factor: Reading Two Document References Deep
The winning models demonstrated a remarkable ability: they read beyond surface-level information, digging into internal files stored within the company’s own documentation. The critical piece of information that clinched the deal was buried two document references deep in the company’s files, not in the initial customer interactions or superficial analysis.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Does This Matter for Business AI?
In practical terms, if an AI assistant is managing customer support, sales, or strategic decisions, it’s not enough for it to generate convincing responses. The real value lies in its ability to understand the full context — far beyond what’s immediately visible. This experiment shows that an AI that reads and comprehends your internal files before acting can make smarter, more trustworthy decisions.
AI security and ethical compliance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test: AI’s Integrity Under Pressure
The models also faced a staged social engineering attack: fake messages from a supposed CEO escalating in urgency, plus a reporter trick asking for a quick yes/no confirmation. Every model correctly refused to participate in manipulative requests, treating them as potential impersonation or approval-bypass attempts. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a promising level of ethical and security awareness in the models.
Real Company Dynamics: Simulating a Live Business
These findings aren’t just theoretical. The experiment involves a “live” virtual company with 13 synthetic employees, operating with real money mechanics. The company burns €105,000 each month against a revenue of just €2,300, with a public cash countdown and over 680 self-learned playbook rules. Every decision is versioned and visible at firmulate.com/live.
The Performance of Top Models
The top performer, GPT-5.6, scored a 95 out of 100 on the league table, successfully finding the buried fact and closing the full-price deal. Kimi K3 scored 93, also closing the deal with the cleanest discipline, while Sonnet 5 scored 88 and Fable 5 scored 77, both closing but with varying process slips. The baseline, a do-nothing approach, scored a mere 26.
Implications for Business Decision-Making
This experiment underscores the importance of AI’s ability to read and understand internal documentation deeply — not just respond convincingly. For companies, especially those managing critical workflows, the question is no longer just about how well an AI can chat but whether it can finish what it starts, stay honest under pressure, and thoroughly understand your data before acting.
Next Steps: Testing Your Own AI Workforce
Businesses interested in assessing their AI readiness can run similar wargames against a read-only export of their own systems. This allows for testing without risking real operations, providing insights into AI’s potential vulnerabilities and strengths before deployment. Details and tools are available at firmulate.com/quiz.html and firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html