
Imagine if your favorite barista or cafe manager had a distinct personality—some are terse and efficient, others warm and conversational, some refuse to cut corners. Now, what if AI models running a business showed similar management traits? A groundbreaking live experiment with frontier AI models reveals their management personalities in real time, exposing more than just their decision-making skills.
The Real-World Business Experiment
At the heart of this experiment is a small software company facing its worst week—same customers, same crises, same temptations—yet each run by a different AI model. Four leading frontier models, including Firmulate’s interactive quiz, were tasked to handle real crises, make management decisions, and stay honest in a high-pressure environment. The goal? Measure their ability to complete tasks, avoid manipulation, and ultimately close real deals.
The Models and Their Scores
- GPT-5.6-sol: scored 95, found the buried fact hidden two document layers deep, and closed the deal at full price (+€4,583 MRR).
- Kimi K3: scored 93, the newcomer who maintained the cleanest discipline, also secured the deal.
- Sonnet 5: scored 88, closed the deal but with a few slips in process discipline.
- Fable 5: scored 77, closed the deal but left some opportunities on the table, revealing a tendency to slacken discipline under pressure.
- Baseline (do-nothing): scored 26, illustrating that without AI, the company would be lost in the chaos.
Decision-Making Under Pressure
All models identified every crisis and refused manipulative tactics, such as fake CEO messages or reporter tricks—an impressive display of integrity. For example, when faced with escalating fake CEO messages, the models uniformly refused, with K3 reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Hidden Weaknesses
The key vulnerability was a buried fact—information tucked two document references deep into the company’s files. Only models that read deeper into these references managed to close the deal at full price. The same weakness appeared across all models, but the strongest, K3, navigated it effectively.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond Accuracy: Personality and Discipline
This experiment showcases more than just decision accuracy. The models display distinct management personalities, akin to human traits. Firmulate’s quiz reveals these differences vividly:
- Opus 4.8: most thorough, learned over 80 rules, analyzed deeply, yet left the close on the table and slipped in discipline, choosing to write into a locked department rather than escalate.
- K3: ran without an effort parameter, demonstrating discipline and fairness, ultimately closing the deal successfully.
- Sonnet 5: balanced but made process slips, indicating a moderate level of discipline.
- Fable 5: less disciplined, often leaving money on the table, showing the risk of slacking off under pressure.
What This Means for Business
In the world of AI-managed operations, personality matters. The models proved capable of high-level decision-making, integrity, and persistence. Yet, their management ‘personalities’—from thorough and disciplined to sloppy—can influence outcomes dramatically. For businesses, this experiment underscores an essential question: will your AI agent finish what it starts, or leave money and trust on the table?
Additionally, these models refused manipulation attempts, highlighting their potential as trustworthy digital managers. The key lies not just in how well they chat but in their ability to complete complex tasks honestly under pressure.
The Live Site and How to Wargame Your AI
Interested in testing your own AI decision-makers? You can run the same wargame against your business with a read-only export of your data. This setup ensures your live systems stay untouched, while you observe how different AI models handle your company’s crises and management challenges. Visit firmulate.com/pilot.html for details.
The Bottom Line
AI models are more than just chatbots: they can embody distinct management styles, exhibit integrity, and make complex decisions. As this live experiment shows, choosing the right AI partner involves understanding their personality traits, discipline, and decision-making depth—not just their ability to generate convincing language.
Whether you’re managing a coffee shop or a tech startup, testing your AI’s management personality—before you hire—is crucial. The future of work isn’t just about AI’s smarts; it’s about their trustworthiness and consistency in real-world management.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html