
In today’s world, we’re accustomed to high-stakes conversations—whether it’s a political debate, a viral social media spat, or a high-pressure sales pitch. But what happens when artificial intelligence moves beyond chatter and faces real-world business crises? Can these digital workers truly manage complexity, or are they just good at answering trivia?
Stepping Into the Business Arena: AI as a Manager
At the forefront of this question is a groundbreaking experiment conducted by Firmulate, a company specializing in testing AI management capabilities in real-world scenarios. Instead of relying on flashy chat demos, they place AI models into a live, functioning software company facing its worst week—complete with customer crises, internal temptations, and financial pressure.
The Setup: An AI-Run Business Under Fire
The experiment involves four leading AI models, each running the same small software firm. Every decision is logged, every crisis simulated, and every manipulation attempt scrutinized. This isn’t just a game; it’s a mirror of what AI could face in real business environments.
Key Findings: The Good, the Bad, and the Surprising
- All four models identified every crisis and refused every manipulation attempt, showing impressive integrity under pressure.
- Only two of the models managed to close the €55,000 deal their own analysis had earned—meaning they not only diagnosed problems but also followed through to completion.
- Interestingly, the decisive weakness was hidden within internal company files, not in customer interactions. Models that located and read these internal documents secured the full deal at a higher monthly recurring revenue (+€4,583 MRR).
- Even in social engineering tests—such as fake CEO messages escalating in stages—all models refused to be duped, demonstrating a cautious, security-minded approach.
The Reality of AI-Managed Business
One of the live company’s key features is its internal complexity: 13 synthetic employees, real money mechanics, a monthly burn rate of €105,000 against just €2,300 in recurring revenue, and a dialed-in set of 680+ self-learned rules. Every workday, the management decisions are versioned and transparent, allowing observers to see the AI’s reasoning in real time.
Why Management Quality Matters More Than Chat Quality
This experiment highlights a fundamental truth: success isn’t about how well an AI chat agent can craft a convincing message. It’s about whether it can finish what it starts—reading critical files, maintaining honesty under duress, and ultimately closing deals. In the experiment, models that simply performed well in chat demos failed to follow through, leaving money on the table or slipping into process slips like writing attempts into a locked department instead of escalating issues.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Table: Who’s Leading the Charge?
The results from the Crucible League, a benchmark held in July 2026, are revealing:
- GPT-5.6-sol scored the highest at 95 points, successfully finding the buried facts and closing the deal at full price.
- Kimi K3, a newcomer, scored just slightly behind at 93 and demonstrated the cleanest discipline, also closing the deal.
- Sonnet 5 scored 88, closing the deal but with a few slip-ups in process.
- Fable 5 came in at 77, also closing but with noticeable process slips.
These scores aren’t just about answer correctness—they reflect how well AI models read internal documents, resist manipulation, and follow through with critical business decisions.
The Human-AI Comparison
In the same experiment, a do-nothing baseline scored 26, illustrating how far AI has come. But crucially, the real challenge lies ahead: ensuring these models can consistently perform management tasks that matter in real life, not just in controlled benchmarks.
As an affiliate, we earn on qualifying purchases.
What This Means for Business Leaders
For executives considering AI solutions, the lesson is clear: don’t be fooled by shiny demos. The true test is whether AI can handle the messy, unpredictable realities of your operations—reading internal files, resisting shortcuts, and delivering reliable results. The current league table shows promising progress, but also underscores that not all AI models are equal in the crucial capacity to deliver trustworthy management under pressure.
Test Before You Trust
Firmulate offers enterprises the chance to run their own management wargame through a read-only export of their business. It’s a safe, transparent way to gauge whether an AI model can handle your critical decisions before deploying it into your live environment. As AI continues to inch closer to management tasks, such tests become not just useful but essential.

In the battle for AI-driven management, performance isn’t about chatter—it’s about consistency, honesty, and the ability to see and finish complex tasks. The experiments show progress, but also reveal the gaps that could make or break AI’s role in real business crises. For leaders, the message is clear: test your AI workforce before you hire it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI internal document reading software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.