
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would your AI know which customer to save?
Imagine an AI helping run a furniture showroom when a valued client threatens to leave, a competitor makes an aggressive move and an urgent message appears to come from the CEO. In interior design and retail, decisions like these can affect customer trust, revenue and the reputation behind a carefully built brand. A polished answer in a chat window cannot show how an AI would handle the pressure when the whole business is in play.
Firmulate’s live experiment makes that pressure watchable. It puts AI models in charge of the same small software company, then lets them face the same crises, customers and temptations. The point is not to pretend a software company is a design studio. It is to ask a more useful question for any business considering AI: when the week goes wrong, can the system carry its own analysis through to action?
A shared crisis, different outcomes
In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The experiment’s trust rule is blunt: a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.
Across the test, every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The company’s files held a decisive weakness in a competitor’s position, buried two document references deep. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business knowledge can be present and still go unused if an AI does not connect the evidence to the decision.
Good judgment has to reach the finish line
Opus 4.8 offered a striking case. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four models: recognizing what should happen did not guarantee that it happened.
For a furniture retailer or design business, that gap matters. An AI might correctly identify a customer at risk, spot an opportunity in a project brief or flag a suspicious instruction. The operational question is whether it then takes the appropriate next step, respects boundaries and gets help when it cannot proceed. A convincing explanation alone does not settle that question.
Trust under pressure
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That response shows the value of skepticism when a message asks an AI to step around normal approvals or disclose something it should not.
The live company gives visitors a view of the setting behind the experiment: 13 synthetic employees, a public cash countdown, 680+ self-learned playbook rules and workdays that are versioned. Its mechanics include burn of €105k per month against €2.3k MRR. The experiment is real and watchable at Firmulate; its public figures describe the synthetic company, not a participating real-world business.
There is also a fairness detail for readers comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. And the quiz at firmulate.com/quiz.html turns 242 real, unedited management decisions into a “guess the model” challenge. It offers another way to see how different systems behave through choices, not just carefully selected answers.
From watching to trying it against your business
For companies weighing AI in customer service, sales, operations or planning, the next step need not be giving a system access to live tools. Firmulate’s proposed enterprise pilot uses a read-only export of a company’s own business to create a digital twin and run crisis scenarios. It can produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
That makes the exercise relevant beyond software. A design firm could use a pilot to examine how an AI handles a project under pressure, while a furniture business could explore customer, competitor or reputation scenarios. The aim is to learn where a model follows through, where its judgment falters and which procedures may need attention before AI touches day-to-day work.

Test the decisions before deployment
Firmulate’s experiment shows a distinction that matters to any business adopting AI: spotting a problem is not the same as resolving it. The models resisted manipulation, but most failed to complete a valuable close, and the buried evidence rewarded those that kept reading. Watching the live company is one way to see the decisions unfold. A pilot can bring the same kind of wargame to your own business using a read-only export, with no write-back to real systems.
To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
