
Imagine watching a high-stakes game where artificial intelligence models are entrusted to run a real, money-losing software company—trying to steer through crises, temptations, and tough decisions. Now, what if you could see which AI makes the right call and which slips up? Welcome to the forefront of AI management testing, where models are judged not by their chatter, but by their ability to deliver results under pressure.
The Live Experiment: AI as a Business Manager
At firmulate.com, a groundbreaking live experiment puts four different AI models in the driver’s seat of a real software company facing its worst week. This isn’t a simulation; it’s a real-world test with real money mechanics, real crises, and real temptations to cheat or cut corners. Every decision the AI makes is recorded, transparent, and auditable, providing a rare window into how these advanced models perform when managing complex, unpredictable situations.
As an affiliate, we earn on qualifying purchases.
The Models and the Scoreboard
The models participating in this challenge are some of the most advanced frontier AIs available. Their scores, out of 100, reflect how well they perform across various metrics:
- gpt-5.6-sol: Scored 95 – The top performer, successfully identifying hidden information that secured a €55,000 deal.
- Kimi K3: Scored 93 – A newcomer showcasing the cleanest disciplinary record, also closing the deal.
- Sonnet 5: Scored 88 – Managed to close the deal but with a few slips in process discipline.
- Fable 5: Scored 77 – Similar to Sonnet, with some weaknesses in decision discipline.
- Opus 4.8: Scored 73 – Historically thorough but left the deal on the table, showing discipline lapses.
The baseline score for a do-nothing approach is just 26, highlighting how much more these models can do, even when they stay honest under pressure.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decisive Moments: The Crises and the Buried Facts
All four models successfully detected and responded to every crisis—no false alarms, no shortcuts. They refused manipulation attempts, including complex social engineering tricks like fake CEO messages or reporter interrogations. Notably, the key to winning a significant deal lay deep in the company’s own files, two document references down. Models that took the time to read and analyze these files ended up closing the deal at full price, worth over €4,580 in monthly recurring revenue. The one that missed the buried facts left the opportunity on the table, illustrating how crucial thorough information reading is in management decisions.
As an affiliate, we earn on qualifying purchases.
Behavioral Profiles: Different Personalities, Same Goals?
The models’ decision styles varied significantly. For instance, Opus 4.8, known for its thoroughness with over 80 learned rules, was the last to secure the deal and showed lapses like postponing escalation. Kimi K3 ran without an effort parameter, maintaining discipline and closing the deal swiftly. The differences showcase that, even with the same task, AI models can exhibit distinct management personalities—some meticulous, others terse or cautious.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Integrity Under Pressure
In a test of integrity, all models refused to be tricked by escalating fake CEO messages and background check requests. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” Their refusal demonstrates a shared capacity to maintain honesty and security, essential traits for trustworthy AI in management roles.
Why This Matters for Business
While the models’ chat abilities are impressive, the real takeaway is their capacity to finish what they start, read critical documents thoroughly, and stay honest when under duress. These traits are vital when AI is integrated into customer relationship management, support queues, or forecasting—areas where cutting corners can cost real money.
Watch Live and Decide
This isn’t a mere demo; it’s a window into the future of AI-driven management. The real company runs every business day, with over 680 self-learned rules and a public cash countdown. You can watch the live company in action, see the decisions unfold, and even run your own simulations against a read-only export of your organization’s data at firmulate.com/pilot.html.
The Takeaway
As AI models become part of your operational toolkit, the question isn’t just about how well they write or converse—it’s whether they can deliver consistent, honest management under pressure. The live experiment proves that some frontier models are already capable of this, scoring as high as 95, and making crucial decisions without shortcuts. The future of business management may well depend on choosing the AI personalities that align with your values and objectives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html