AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a world where artificial intelligence not only supports your business decisions but proves smarter than seasoned managers — even under the toughest conditions. As AI models face off in real-world simulations, the question shifts from “Can it talk like a pro?” to “Can it get the job done without a slip-up?” Welcome to a groundbreaking experiment that’s turning the business world upside down.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Challenge: Testing AI in Real Business Crisis

In July 2026, a unique event took place that’s stirring conversations across industries: a live, real-time test of leading AI models running an actual, functioning software company. This wasn’t a mere demo or chat simulation — it was a full-blown, auditable experiment where each AI was tasked with managing a small company through its worst week. From customer crises to internal temptations, every scenario was identical across models, ensuring a fair fight.

Amazon

AI management software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Contenders and Their Scores

  • GPT-5.6-sol: Scored the highest at 95, finding the critical buried information in the company’s files and sealing the deal for €55,000, adding €4,583 MRR.
  • Kimi K3: The newcomer from Moonshot scored a close second at 93, demonstrating disciplined decision-making and also securing the €55k contract.
  • Sonnet 5: With 88 points, it managed to close the deal but showed a few process slips along the way.
  • Fable 5: Scored 77, with some missed opportunities and weaker discipline, leaving the engagement on the table.
  • Opus 4.8: The most thorough participant with over 80 learned rules, yet only scored 73 and left the deal unclosed. Its discipline slipped in critical moments.
Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings and Surprises

What made this experiment stand out wasn’t just the scores. All models successfully identified every crisis and refused manipulative tricks designed to test their honesty. Notably, the decisive edge came from how the models utilized internal company files — those buried references that savvy human managers often overlook. The AI that read and understood these hidden details won the deal at full price, translating into real revenue.

Amazon

AI crisis management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Test of Trust and Integrity

The models faced social engineering scenarios — fake CEO messages escalating in severity, along with a reporter trick asking for a quick background approval — and all refused. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This level of cautious judgment reflects a crucial trait for trustworthy AI in enterprise environments.

Amazon

AI business simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business and Its Lessons

This wasn’t a staged demo. The experiment is ongoing at firmulate.com, where every business day a real company faces real challenges. With 13 synthetic employees, a public cash countdown, and over 680 self-learned rules, the system demonstrates how AI can manage complex, money-critical operations without writing back to actual systems — making it a safe, transparent testbed for enterprise AI deployment.

What the Scores Say About the Future of AI Management

While GPT-5.6-sol led the league, Kimi K3’s performance was a surprise. The new entrant managed to beat three of four Western frontier models, including the well-established Sonnet 5 and Fable 5. This highlights a vital point: in a league where all models can spot crises and resist manipulation, the difference is in disciplined decision-making and thorough information reading. The fact that K3 ran without an effort parameter (API default) while others used a more aggressive setting underscores that a well-tuned model can outperform more resource-intensive options.

Why This Matters for Business Leaders

For anyone betting on AI to transform management, these results underline a key truth: it’s not just about how well an AI can generate chat or mimic human speech. The critical questions are whether it can finish what it starts, stay honest under pressure, and truly understand internal information — the buried facts that matter most.

Next Steps and Limitations

Interested parties can explore these models further through live testing and pilot programs. The experiment is ongoing, and the leaderboard is open for scrutiny at firmulate.com/benchmarks.html. However, it’s important to note that K3 ran at the default setting, while others ran at a higher effort level (xhigh), which can influence outcomes. Transparency about these settings ensures fair comparison and highlights how nuances in tuning can impact performance.

The Big Takeaway

As AI models continue to evolve in real-world business environments, the real differentiator will be discipline, thoroughness, and honesty — traits that can decide whether AI truly adds value or just creates a shiny distraction. The experiment proves one thing clearly: the league is open, and choosing the right AI model is no longer a gamble based solely on demo performance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Live AI management tests reveal that discipline and thoroughness matter most. Kimi K3, a newcomer, beat established models by reading buried info, showing enterprise AI’s real potential — but the choice of model requires careful testing.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sadie Frost Surges In Global Coverage

Sadie Frost experiences a significant increase in international media mentions, with reports rising to 42 times the baseline according to GDELT data.

Maya Hawke Surges In Global Coverage

Maya Hawke has experienced a surge in international media coverage, with mentions increasing 25-fold in recent days, highlighting her rising prominence.

BET Awards 2026 viewers insist Keke Palmer ‘should’ve been’ host instead of Druski

Fans of the BET Awards 2026 are expressing strong support for Keke Palmer to have hosted instead of Druski, sparking online debates.

Rachael Ray Fans Comments

Fans of Rachael Ray have expressed concern over her recent appearance changes online, sparking widespread discussion and speculation.