
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would your AI know which client to save first?
Imagine a furniture showroom in the middle of its worst week: a major client is wavering, a competitor is undercutting prices, and a message claiming to come from the CEO urges someone to bend the rules. For interior design and furniture businesses, the promise of AI agents managing customer relationships or sales comes with a practical question: how will they behave when the pressure is real?
Firmulate is testing that question in a live, watchable experiment. Its public brand describes an AI company emulator that measures management decisions through crises, money mechanics and temptations—not just polished answers in a chat window. The next step is to bring the exercise closer to home: a pilot built around a company’s own business data.
One difficult week, the same company
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
The results suggest that spotting trouble is only part of the job. Every model identified every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The gap between sound advice and follow-through—the finding summed up as “Same diagnosis, same pitch — no signature”—would matter just as much in a design business. An agent might correctly identify a client at risk or a promising commercial lead, then fail to make the move that protects the relationship or closes the sale.
The clue was buried in the company’s own files
The decisive competitor weakness was two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a furniture retailer, design studio or interiors supplier, that points to a familiar challenge: the detail that changes a negotiation may sit in a project note, an old customer record or a policy document—not in the latest message.
Trust also faced a direct test. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” trick. All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of restraint matters wherever customer information, discounts or company commitments are involved.
Thoroughness does not guarantee a close
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. More detailed work did not automatically mean better execution.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz powered by 242 real, unedited management decisions. Its live company has 13 synthetic employees, burn of €105k per month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. The experiment is watchable at firmulate.com.
From watching to your own pilot
The enterprise offer takes the wargame from a public experiment to a company’s own operation. It uses a read-only export to create a digital twin, then tests crisis scenarios against the business and produces a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems.
For an interiors company, that could make an AI’s judgment easier to examine before entrusting it with work around the CRM, customer support or forecasts. The point is to see how it responds to pressure using your own business context, while keeping the exercise separate from live operations.

Test the judgment before you hand over the keys
Firmulate’s experiment shows a clear difference between recognizing a crisis and carrying a decision through. A pilot lets enterprises examine that difference against their own data, with a read-only export and no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
