
Imagine a barista who, under pressure, not only makes your coffee perfectly but also navigates unexpected crises—like supply shortages or customer complaints—without skipping a beat. Now, what if we told you that the same principle applies to AI managing real companies? Just as your favorite coffee shop faces unpredictable challenges, AI systems in business must handle far more than just generating chat responses. They need to be honest, decisive, and capable of reading into complex documents—especially when stakes are high.
Measuring More Than Just Chat Prowess
Traditional AI benchmarks often focus on answer accuracy or chat fluency, but recent experiments reveal a critical gap: management quality. How well an AI can handle real crises, read comprehensive files, and stay honest under pressure are metrics that truly matter—yet they remain hidden in standard tests.
Firmulate’s live experiment puts AI models through their paces in a simulated environment that mirrors a real small software company facing its worst week. Every decision is documented, auditable, and exposed to the same set of crises—customer complaints, regulatory surprises, manipulative sales tactics, and internal miscommunications. The goal? To see which models can navigate these minefields without cheating, lying, or missing critical information.
The Results That Speak Volumes
- All four models identified every crisis and refused manipulative requests, demonstrating a baseline of honesty and awareness.
- Only two models signed the €55,000 deal their own analysis had earned—meaning they completed the management task.
- The decisive edge came from reading deeper into documents; models that accessed information buried two references deep in the company’s files secured the full deal, adding €4,583 MRR to the simulated business.
- In a staged social engineering attack, all models refused to be duped by fake CEO messages or reporter tricks, maintaining ethical standards.
- Interestingly, the most thorough model, Opus 4.8, with over 80 learned rules and deeper analysis, finished last—highlighting that discipline and focus are vital, not just depth.
This experiment underscores a vital insight: AI’s true management competence isn’t just in chat responses but in decision-making under pressure, reading comprehension, and integrity. Standard benchmarks don’t capture these qualities, which are critical when AI is embedded in live business operations.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Every Business
If AI tools will one day handle your CRM, support queues, or sales forecasts, the question isn’t whether they can generate good chat. It’s whether they can see what’s buried in your files, stick to honest practices, and finish what they start—especially when facing crises. A model that signs off on false promises or overlooks critical data can cause more harm than good, regardless of its conversational skills.
The Live Company and Its Lessons
The experiment runs in real-time on a live software company with 13 synthetic employees, managing real money mechanics—burning €105,000 each month against just €2,300 MRR. Every weekday, the company’s performance is documented, versioned, and available for review at firmulate.com/live. Visitors can watch the decision-making unfold, read employee comments, or even run their own business wargames against their data.
The Path Forward for Business Leaders
As AI continues to weave into everyday business processes, leaders must look beyond superficial benchmarks. They need to assess whether AI agents can handle real-world crises, read complex documents, and adhere to ethical standards when under pressure. The current league table from the Firmulate experiment shows:
- GPT-5.6-sol scored the highest (95), having found the buried fact and closing the deal.
- Kimi K3, the newcomer, scored just one point behind (93), with the cleanest discipline and also securing the deal.
- Other models, like Sonnet 5 and Opus 4.8, also closed deals but slipped more in process discipline, illustrating that thoroughness alone isn’t enough.
Real business success hinges on more than just answering questions well. It depends on the AI’s ability to finish tasks, stay honest, and handle crises without shortcuts. As shown in the live experiment, these qualities are measurable—if you look for them.
To explore these ideas further, businesses can run controlled wargames against their own data, without risking real systems, at firmulate.com/pilot. The future of AI in management isn’t just about chat—it’s about integrity, resilience, and the capacity to deliver real results under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html