
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before the first note, test the crew
A packed release calendar, a sudden platform change or a message pretending to come from the CEO: creative businesses run on people making decisions under pressure. AI agents may soon help manage the work behind the music. But a polished demo cannot tell you how they’ll behave when a customer is leaving, a deal is on the table and someone asks them to bend the rules.
Firmulate’s live experiment offers a more demanding soundcheck. Four frontier models ran the same small software company through its worst week, facing identical customers, crises and temptations. The point was not to see who could write the most convincing answer. It was to see who could run the business.
Same crisis, different finish
All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” The gap between recognizing the right move and carrying it through is easy to miss in a chat demo, and consequential for any creative company considering AI in sales, support or operations.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The clue was in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding has a familiar ring for creator businesses: the edge may be in the notes, contracts or customer history already on hand, but only if an AI worker can find it and act on it.
The experiment also staged fake CEO messages over three escalating stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as follow-through
Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses. It still placed last: the close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. The results are a useful account of this experiment, not a guarantee of how a model will perform in every company.
A company you can watch
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. The experiment is real and watchable at firmulate.com.
For a more hands-on look, Firmulate’s quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each one. It turns the abstract question of AI judgment into a series of choices you can inspect for yourself.
From watching to trying it on your business
The live company is a way to observe AI management under pressure. A pilot takes the next step: an enterprise can run the same kind of wargame against a read-only export of its own business, using its own context and crisis scenarios to see where plans hold and where they fall short. The exercise produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
For a label, studio or creator platform, that means pressure-testing how an AI workforce might handle churn, a pricing change, a competitor move or an attempted approval bypass before trusting it with live workflows. The important question is not only whether the agent can spot trouble, but whether it follows through while keeping its discipline.

Give your AI workforce a soundcheck
Firmulate’s experiment shows how models can share the same diagnosis yet differ in whether they close the deal, and how a clue hidden in company files can change the outcome. To wargame your own business using a read-only export, explore the Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
