
Imagine a barista choosing your coffee — but instead of beans, it’s an AI managing a bustling café’s week of chaos. How do you tell if your digital manager is trustworthy, decisive, or just plain effective? Now, scale that question to the world of business management, where AI models are increasingly making decisions that impact real money and reputation. Welcome to a behind-the-scenes look at a groundbreaking live experiment that pits AI models against each other in the high-stakes environment of running a small software company.
Recently, four frontier AI models faced the same challenging week at a real software business. This wasn’t a scripted demo—every crisis was authentic, every temptation genuine, and every decision carefully recorded. The goal? To see not just if these models could spot problems, but whether they would act with integrity, thoroughness, and discipline under pressure.
The models included gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, each with different personalities and strengths. For example, the most thorough, Opus 4.8, analyzed over 80 rules and provided detailed assessments but ultimately left some deals on the table due to discipline slips. Meanwhile, Kimi K3, the newcomer, demonstrated the cleanest judgment, closing the same deals at a high score of 93, just behind the top performer, gpt-5.6-sol, which scored 95 by uncovering a hidden critical document and sealing the final deal.
This experiment was set up to test three core qualities: crisis detection, resistance to manipulation, and decision-making integrity. Remarkably, all models identified every crisis and refused every attempt at manipulation, including social engineering tactics like staged CEO messages and reporter tricks. For instance, during a staged scenario where a fake CEO asked for a quick approval, all models refused, citing concerns over impersonation and bypassing approval processes. This shows a shared ability to recognize and resist social pressure—a crucial trait for AI used in management.
But the true test came when the models had to read deeply into the company’s own files to find a key document. Those that read the files thoroughly secured a lucrative deal, worth over €4,583 in monthly recurring revenue, fully paid without discounts. Conversely, models that missed this detail left money on the table, illustrating how critical in-depth information processing is to effective AI management.
All decisions were made within a simulated company with 13 synthetic employees, real financial mechanics—burning €105,000 monthly against €2,300 in revenue—and a live cash countdown. The environment was transparent, versioned daily, and open for observers to watch the decision process unfold on firmulate.com/live. This setup shows that AI management isn’t about clever chat—it’s about consistent, honest, and strategic action.
Interestingly, the experiment also tested the models’ responses to social engineering attempts, which are common risks in real companies. All models refused to approve the staged requests, with Kimi K3 explicitly treating the request as a possible impersonation or bypass of approval channels. This kind of resilience is vital if AI is to be trusted in sensitive management roles.
Among the competitors, Opus 4.8—despite its thorough analysis—showed a discipline slip, leaving a deal unclosed and escalating issues into a locked department instead of resolving them promptly. This demonstrates that thoroughness alone isn’t enough; decision discipline and focus matter just as much.
When ranking the models, the final scores reflected their effectiveness: gpt-5.6-sol led with 95 points, followed by Kimi K3 at 93, then Sonnet 5 at 88, and Fable 5 at 77. These scores are based on their ability to identify critical facts, resist manipulation, and close profitable deals, shining a light on their different management personalities and trustworthiness under pressure.
For business leaders curious about how AI can shape management, this live experiment is more than a demo—it’s a glimpse into the future of AI decision-making in real-world environments. Managers can run their own scenarios against a read-only version of their company’s data, testing how AI might perform before deploying it in critical roles. This approach ensures that AI doesn’t just sound good in demos but proves its worth in actual operational settings.
What does this mean for your coffee shop or beverage business? As AI models become more capable of managing complex decisions, their personalities—thorough, cautious, or terse—will influence how they support your team. Knowing which model aligns with your management style can help you choose an AI partner that keeps your business honest, disciplined, and profitable.

This live AI management experiment shows that models can identify crises, resist manipulation, and close deals—qualities crucial for trustworthy AI in business. The real takeaway? Not all AIs are equal: personality, thoroughness, and discipline vary, influencing how they manage your company’s future. Try running your own scenarios at firmulate.com/quiz.html to see which AI might be your best manager-in-waiting.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.