
This isn’t science fiction — it’s a live experiment in AI-driven management
Imagine a company entirely run by artificial intelligence, with no human staff, battling everyday crises while losing money every month. Now, watch it unfold live — a real, auditable experiment that reveals whether AI can actually manage a business under pressure. For those in the coffee, tea, and beverage industry, where managing supply chains, customer relationships, and quality control are daily battles, this story offers a glimpse into the future of automation and decision-making.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company: An AI-Driven Business in Action
At the heart of this experiment is a small software company, operating with 13 synthetic employees and real money mechanics. Every workday, its decisions are versioned, audited, and made in a transparent environment accessible at firmulate.com/live.html. The company spends €105,000 a month but earns just €2,300 in recurring revenue, illustrating the stark reality of a business in survival mode. It’s a literal battleground for AI decision-making, where every crisis — from customer issues to internal conflicts — is simulated and monitored.
The Scoring and the Frontiers
Four leading AI models competed by running the same worst-week scenario, facing the same crises and temptations. Their scores ranged from 77 to 95, with the top performer, gpt-5.6-sol, successfully uncovering hidden data that clinched a €55,000 deal. Meanwhile, other models struggled with discipline or overlooked critical files, leading to missed profit opportunities.
Can AI Win Deals and Maintain Integrity?
In a critical test, all models faced social engineering — fake CEO messages and a reporter’s subtle pressure. Every model refused to be manipulated, demonstrating a fundamental understanding of trust and risk. Yet, only two models signed the deal based on their own analysis; one with a perfect score and the other with a slightly lower one. The decisive factor? Reading and understanding buried internal documents, not just surface-level customer data.
Management and Decision-Making in the Wild
This experiment measures management quality, not just chatbot chatter. The most thorough participant, Opus 4.8, with over 80 learned rules, performed the worst, often leaving deals on the table and slipping into internal silos instead of escalating issues. Contrastingly, other models showed greater discipline or better decision patterns, but no AI was infallible.
What This Means for Your Business
For industries like beverage production and retail, where operational decisions are critical, the question isn’t whether AI can generate convincing dialogue. It’s whether AI can finish tasks, read the right information, stay honest under pressure, and produce useful work. This experiment offers a rare, transparent look at how AI might replace or support human management in complex, real-world settings.

Key Takeaways
This live experiment shows that AI can recognize crises and resist manipulation but still struggles with full deal closure and disciplined management. The strongest model uncovered hidden internal data that led to a major deal, proving that reading and understanding internal documents is crucial. As AI tools become more integrated into business operations, the questions for industry leaders are simple: Will your AI work under pressure? Will it stay honest? And will it get the work done — profitably and reliably?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html