
Imagine hiring an AI to manage your client relationships or handle complex decision-making. You might think its responses depend solely on how well it chats or summarizes. But in a groundbreaking live experiment, the real test was whether these AIs could read deeply into company files—two layers deep—before making a decision. The results reveal a crucial, often overlooked property that could determine whether your AI partner is worth the investment.
The AI Experiment That Revealed a Hidden Weakness
Recently, four state-of-the-art AI models faced a simulated crisis within a small software company. Their task? Navigate a week filled with customer issues, economic pressure, and manipulative tactics—all in a controlled, real-time environment. The goal was simple but vital: see if the AIs could identify critical hidden information buried in the company’s own files, not just react to surface-level customer interactions.
The experiment was designed to mimic real-world corporate decision-making, complete with record-keeping, file analysis, and trust challenges. Each AI was tasked with diagnosing issues, resisting manipulations, and ultimately closing a significant deal worth €55,000 in monthly recurring revenue. Every decision was recorded and auditable, ensuring transparency.
The Key Findings
All four models successfully detected every crisis and refused manipulative attempts, demonstrating their ability to uphold ethical standards under pressure. However, only two managed to close the deal—those that read the company’s files thoroughly and understood the deeper context. The other two, despite making the right diagnosis and giving the correct pitch, left the deal on the table.
The decisive factor lay two document references deep in the company’s internal files. Models that took the time to read and analyze these buried facts earned the full €55,000 deal, translating into an additional €4,583 monthly recurring revenue. Conversely, models that overlooked these crucial details failed to secure the agreement, even when their surface analysis was correct.
The Significance for Business and AI Adoption
This experiment underscores an essential property for AI agents in corporate environments: the ability to read and interpret a company’s internal documentation before making decisions. It’s not enough for an AI to respond well in chat; it must access and understand the full depth of relevant data. This ability to perform multi-hop reasoning—reading documents, synthesizing information, and making informed choices—is a measurable, critical factor influencing revenue outcomes.
As an affiliate, we earn on qualifying purchases.
How AI Handles Manipulation and Trust
The experiment also tested the models against social engineering tactics—fake CEO messages escalating in stages and an attempt to trick the AI with a reporter’s background query. Remarkably, all five models refused to be manipulated. Kimi K3 explained its reasoning explicitly: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a vital aspect of AI trustworthiness, especially in sensitive decision-making contexts.
The Real-World Application: Firmulate’s Live Business Simulation
Beyond the lab, this research is embodied in a live, observable setup. The company’s simulation includes 13 synthetic employees managing real-money mechanics—burning €105k monthly against a €2.3k monthly recurring revenue. Every day, the AI operates within a complex environment aligned with real business goals, and the entire process is visible at firmulate.com/live. This transparency allows organizations to evaluate their AI’s ability to make trustworthy, informed decisions before deploying it in their actual operations.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business Decision-Making and AI Deployment
For interior designers, furniture retailers, or decor specialists, the takeaway may seem distant. But the truth is, any industry adopting AI for customer engagement, project management, or sales will face similar challenges. It’s not just about how well an AI writes or summarizes; it’s whether it can finish what it starts—reading the necessary background, resisting manipulation, and making decisions based on complete understanding.
The benchmark scores reinforce this point: the top model, GPT-5.6-sol, scored 95 out of 100, successfully closing the deal by uncovering hidden facts. Kimi K3, though a newcomer, scored 93 and demonstrated the cleanest discipline. Other models, despite their competence, left crucial information unexamined and missed revenue opportunities.
multi-hop reasoning AI solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Future of AI in Business
As AI continues to mature, businesses must prioritize models that excel in thoroughness and trustworthiness. The experiment shows that even the most advanced models can falter if they overlook the deep, buried information essential for decisive action. For companies, the question isn’t just about AI’s ability to generate appealing responses but whether it can perform the complex reasoning necessary to deliver real value and secure deals.
Organizations interested in testing their own AI workforce can run similar simulations via dedicated tools, ensuring their agents are prepared to handle real crises without writing back to their systems—safeguarding integrity while optimizing performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI file reading and comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.