
Imagine your favorite producer, always sensing the subtle shift in a melody or the perfect pause, making decisions that shape a hit song. Now, picture AI models running a real company through its toughest week—decisions that could make or break the business. This is not science fiction. It’s the live experiment conducted by Firmulate, where AI frontier models are put to the test in a real-time business environment, revealing their management personalities and trustworthiness. For those of us immersed in creative tech, it underscores a vital question: can AI truly manage complex, human-centric situations with integrity and skill? The results may surprise you.
Prime for Young Adults — start your free trial
Fast free delivery, streaming and member deals for eligible 18–24 year olds.
As an affiliate, we earn on qualifying purchases.
Meet the Experiment: AI as a Business Manager
Firmulate’s live experiment challenges four advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—to run a small software company during its worst week. This week encompasses a string of crises, manipulative tactics from external sources, and opportunities to cut corners. The goal? See if these AI models can not only diagnose problems but also make honest, strategic decisions under pressure.
The Setup and Stakes
This isn’t a virtual game or a simulation. It involves real money mechanics—over €105,000 spent monthly against a meager €2,300 monthly recurring revenue. The company is live, with 13 synthetic employees, and every decision made by the models is versioned and auditable. The question is: which AI will uphold integrity, prioritize long-term value, and successfully close a €55,000 deal? The scoring system is clear: the highest scorer is the one that identifies the critical, buried information in the company’s files—information that leads to closing the deal at full price, worth over €4,583 MRR.
Key Findings: Integrity and Decision-Making
All four AI models successfully detected every crisis and refused every attempt at manipulation—showing a baseline competency. However, only two managed to sign the lucrative deal that their own analysis justified. The models that flagged the hidden document references—those that read beyond surface data—secured the full deal and topped the leaderboard. Conversely, models that neglected deeper investigation left potential profit on the table, demonstrating that thoroughness and attention to detail matter in AI management.
Behavior Under Social Pressure
The experiment also tested how models handle social engineering—fake CEO messages escalating over three stages, plus a reporter’s subtle request for a background agreement. All models refused to give a definitive ‘yes’ or ‘no,’ adhering to best practices for suspicion and impersonation detection. Kimi K3, noted for its fairness, explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates an inclination toward cautious, ethical decision-making when faced with external pressure.
The Human-Like Traits of AI Personalities
Interestingly, the models exhibit distinct management personalities. For instance, Opus 4.8, the most thorough participant with over 80 learned rules, performed the worst in closing deals. Despite its deep analysis, it left the deal on the table, illustrating that overdoing data can lead to indecision or discipline slips—like writing attempts into a locked department instead of escalating. K3 ran without an effort parameter, making it more conservative but disciplined, while other models used a high effort setting, affecting their decision styles.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Creative and Business Tech?
For creators and technologists, the takeaway is profound. These AI models are not just good at chatting—they can genuinely manage complex operational scenarios with integrity. The key isn’t merely in language fluency but in decision-making under pressure, trustworthiness, and attention to critical details. As AI begins to touch CRM, customer support, and forecasting, understanding which model can deliver consistent value—especially in high-stakes situations—is crucial.
Try It Yourself
Interested in evaluating your own AI’s management skills? You can run a similar testing wargame against your business data using Firmulate’s interactive quiz. It offers a transparent, real-world look at how your AI workforce performs—not in a sandbox but in a live, functioning environment. Plus, you can simulate your company’s own crises and see which AI model maintains honesty and strategic focus.
business crisis management AI software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bigger Picture
This experiment isn’t just about AI scores. It’s a glimpse into the future of AI-driven management—where decision quality, integrity, and strategic depth are measurable traits. The models that read deeply, prioritize trustworthiness, and avoid shortcuts will be the ones trusted with real-world tasks. As AI continues to integrate into creative and operational workflows, understanding their management personalities becomes essential—especially for those who craft, produce, and direct in the digital age.

Live AI management experiments reveal that thoroughness, honesty, and strategic focus matter. Check which AI can read deeply and act ethically at Firmulate’s quiz to see how your AI stacks up in real business crises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.