
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Good output is only half the job
A music creator’s AI assistant can draft a release plan, sort feedback or pitch a brand partnership. But can it read the fine print, spot the real opportunity and carry the work through to a signed deal? Firmulate put that broader question to frontier models by asking them to run a software company through its worst week.
A company under pressure
The experiment gave each model the same customers, crises and temptations, with decisions versioned and auditable. Firmulate’s synthetic company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, while a public countdown tracks its cash. The company runs every business day, and its playbook has accumulated more than 680 learned rules. The live experiment is watchable at Firmulate.
The results complicate the usual question of which model writes best. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. In the final July 2026 league, gpt-5.6-sol scored 95, Moonshot’s Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total.
The detail that changed the outcome
The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That distinction has an obvious echo for creators: a capable assistant may summarize the brief, but useful work can depend on checking the contracts, campaign notes or past audience feedback around it.
Kimi K3 finished second overall, ahead of three of the four Western frontier models in the comparison. It found the buried security needle, won the deal, saved the churning customer and resisted all three baits, with one deviation—the cleanest discipline in the field. One of those tests used a fake CEO message that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The comparison also exposed a gap between analysis and action. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Firmulate’s benchmark results frame the finding plainly: good diagnosis does not guarantee follow-through.
Why a creator should care
For an independent artist, label or production studio, an AI tool’s polished response is easy to notice. Its handling of a sensitive request, its willingness to check source material and its ability to finish a task are harder to judge from a demo. Those differences matter when agents touch audience data, support messages, budgets or business forecasts.
Firmulate says 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can also run the same wargame against a read-only export of their own business; the pilot does not write back to real systems.

Test the work, not just the voice
Kimi K3’s result makes the model race look more open, but the narrow gap between first and second is not a universal buying guide. Picking a model without testing it on your own work is a bet. For creator businesses, that means checking whether an assistant can protect trust, find the detail buried in your files and finish the job—not only whether it sounds convincing.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
