
Imagine observing a company that has no human employees, yet faces the same crises, temptations, and tough decisions as any traditional business. Now, imagine watching this company burn through €105,000 every month, while generating just €2,300 in recurring revenue. Welcome to the world of Firmulate, where AI models run a live company in front of your eyes, revealing what it really takes to build trustworthy, resilient automation.
Inside the Live-Run AI Company
Firmulate operates a unique experiment: a small software business managed entirely by AI models, each acting as a virtual employee. This company, with 13 synthetic team members, is not just a simulation—it’s a real-time, transparent battle for survival, with daily updates and every decision logged and versioned. The company’s cash is rapidly depleting—burning €105,000 a month—against a modest €2,300 monthly recurring revenue. The digital storefront is visible at firmulate.com/live, inviting anyone to observe the chaos and discipline of AI-driven management.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment and Its Findings
The core idea was simple yet profound: subject four leading AI models to the same grueling week of decision-making, crises, and ethical dilemmas, all under the same conditions. These models included GPT-5.6, Kimi K3, Sonnet 5, and Fable 5. Each was tasked with navigating this turbulent simulated environment, making decisions that affected the company’s fate.
Remarkably, all four models identified every crisis, from customer churn to internal mishaps, and refused every manipulation attempt—be it fake CEO messages or subtle bribery offers. Nevertheless, only two of them managed to close the €55,000 deal their own analysis had earned. The other two, despite accurate diagnoses, failed to follow through—highlighting a crucial gap between recognizing opportunities and executing them.

The AI Cybersecurity Handbook
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses in AI Decision-Making
The decisive weakness was not in the models’ ability to spot crises but in their understanding of internal documents. In fact, vital information buried two document references deep in the company’s files was the key to sealing the deal. Models that read and interpret these shared files gained the upper hand, closing the deal at the full price of over €4,580 in monthly recurring revenue.

I, Human: AI, Automation, and the Quest to Reclaim What Makes Us Unique
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Ethical Resilience Under Pressure
Throughout the experiment, social engineering attempts were met with unwavering refusal. Fake CEO messages escalating over three stages, even a reporter trick asking for a suspicious ‘yes/no’ on background—every model refused. Kimi K3, noted for its fairness, explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that trustworthiness isn’t just coded in; it’s actively upheld even in simulated high-pressure scenarios.

Using AI at Work: Time Management for Busy Professionals: A Non-Technical, Tool-Agnostic Playbook to Prioritize Better, Control Your Calendar, and … Week (Leadership Coaching by Jess Pryce 9)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications for Business Automation
What does this reveal for businesses considering AI automation? The experiment shows that success isn’t just about AI writing convincingly or managing customer interactions. It’s about the AI’s capacity to finish what it begins, to read all relevant information, and to resist manipulation when stakes are high. A model may diagnose an issue perfectly but still falter in executing a deal or following through on promises, especially when discipline slips or temptations grow.
Discipline, Rules, and Building Trust
The company’s operational discipline was measured through over 680 self-learned playbook rules, with every decision versioned and auditable. Notably, Opus 4.8, the most thorough participant with over 80 learned rules, finished last—its discipline slipping and an unexecuted deal leaving a critical opportunity on the table. This underscores an essential point: in complex decision environments, thorough rules and disciplined behavior are critical, but so is the ability to adapt and escalate when necessary.
Why Interior Design and Business Trust Have More in Common Than You Think
While this experiment might seem distant from interior design or furniture, the core lesson resonates: trust, discipline, and the ability to follow through are universal. Just as a well-designed space must function smoothly under pressure, a trustworthy AI-driven company must perform reliably under stress. Transparency, thoroughness, and unwavering decision-making become the foundation for building confidence—be it in a living room or a business process.
The Future of Honest AI in Business
As this experiment continues to unfold, one thing is clear: AI models can identify crises, refuse manipulation, and even close deals—if they are built with discipline and integrity at their core. The question isn’t just whether an AI can produce convincing language; it’s whether it can truly deliver honest, consistent results over time. The live experiment at firmulate.com/live makes this question tangible and urgent.
What You Can Do Today
For businesses contemplating automation, run your own wargame against your real processes. Using Firmulate’s platform, you can deploy a read-only export of your operations and see how your AI workforce handles crises, temptations, and the critical moments that define success. The future belongs to those who understand that building trust and discipline into AI is not optional—it’s essential.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html