firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Good output is only half the job

A music creator’s AI assistant can draft a release plan, sort feedback or pitch a brand partnership. But can it read the fine print, spot the real opportunity and carry the work through to a signed deal? Firmulate put that broader question to frontier models by asking them to run a software company through its worst week.

A company under pressure

The experiment gave each model the same customers, crises and temptations, with decisions versioned and auditable. Firmulate’s synthetic company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, while a public countdown tracks its cash. The company runs every business day, and its playbook has accumulated more than 680 learned rules. The live experiment is watchable at Firmulate.

The results complicate the usual question of which model writes best. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. In the final July 2026 league, gpt-5.6-sol scored 95, Moonshot’s Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total.

The detail that changed the outcome

The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That distinction has an obvious echo for creators: a capable assistant may summarize the brief, but useful work can depend on checking the contracts, campaign notes or past audience feedback around it.

Kimi K3 finished second overall, ahead of three of the four Western frontier models in the comparison. It found the buried security needle, won the deal, saved the churning customer and resisted all three baits, with one deviation—the cleanest discipline in the field. One of those tests used a fake CEO message that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The comparison also exposed a gap between analysis and action. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Firmulate’s benchmark results frame the finding plainly: good diagnosis does not guarantee follow-through.

Why a creator should care

For an independent artist, label or production studio, an AI tool’s polished response is easy to notice. Its handling of a sensitive request, its willingness to check source material and its ability to finish a task are harder to judge from a demo. Those differences matter when agents touch audience data, support messages, budgets or business forecasts.

Firmulate says 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can also run the same wargame against a read-only export of their own business; the pilot does not write back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work, not just the voice

Kimi K3’s result makes the model race look more open, but the narrow gap between first and second is not a universal buying guide. Picking a model without testing it on your own work is a bet. For creator businesses, that means checking whether an assistant can protect trust, find the detail buried in your files and finish the job—not only whether it sounds convincing.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Benchmarks Reveal the Truth About Management Skills — Not Just Chatting

Firmulate’s live benchmark tests AI models managing a mock company through crises, revealing strengths in honesty, discipline, and operational performance—baving the way for smarter, trustworthy AI deployment.

The Role of Music in Art Therapy

Get ready to uncover how music enhances emotional expression in art therapy, unlocking pathways to healing and self-discovery that you never knew existed.

Prince Harry Surges In Global Coverage

Prince Harry’s coverage surges worldwide, with GDELT reporting 32 mentions in recent days, marking a notable increase in media focus.

Social Media Regulation: Protecting Kids Online in 2025

Just as social media evolves, new regulations in 2025 aim to better protect kids online—discover how these changes could impact their digital safety.