
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What a Garden Center Can Teach Us About AI’s Business Potential
Imagine managing a bustling greenhouse during a week of unexpected crises: supply chain hiccups, customer disputes, and tricky decisions—all while maintaining honesty and composure. Now, picture AI models navigating the same chaos, revealing not just what they can say but what they can truly do. That’s the heart of a groundbreaking real-world experiment happening now, where AI isn’t just chatting — it’s running a real company in real time.
As an affiliate, we earn on qualifying purchases.
Inside the AI Company Emulator: Real Decisions, Real Stakes
Firmulate’s live experiment puts four leading AI models through their paces by running a small software company facing its worst week. The goal isn’t just to generate convincing conversation but to actually manage crises, read critical internal files, and stay honest under pressure. Every decision is tracked and auditable, mimicking the complexities of real business management.
The League of AI Performers
- gpt-5.6-sol: scored the highest with 95 out of 100, finding buried information, closing a key deal, and demonstrating full performance.
- Kimi K3: the newcomer from Moonshot, scored 93, just behind the leader, and was praised for its discipline and accuracy.
- Sonnet 5: scored 88, managed to close the deal but showed some process slips.
- Fable 5: scored 77, also closing the deal but with noticeable weaknesses in process discipline.
Remarkably, all models identified every crisis and refused manipulative attempts—like fake CEO messages or reporter tricks. Yet, only two actually signed the €55,000 deal they had analyzed and recommended, showing that performance isn’t just about recognizing problems but acting decisively and ethically.
Key Discoveries and Surprising Weaknesses
The decisive weakness was not in reading customer emails or reports—it lay two document references deep in internal files. The models that read these files secured the deal at full price, adding €4,583 MRR to the company’s bottom line. Meanwhile, models that failed to look beneath the surface left money on the table, highlighting that thorough internal knowledge is crucial for real-world AI management.
Staying Honest Under Pressure
Social engineering attempts, like staged CEO approvals and staged reporter questions, met unanimous refusals. Kimi K3’s reasoning was clear: treat suspicious requests as possible impersonation attempts. This discipline is vital in safeguarding against manipulation, especially as AI begins to handle more sensitive company functions.
The Company Behind the Experiment
The live company used for testing comprises 13 synthetic employees with real money mechanics—burning €105k/month against €2.3k MRR. It’s a microcosm of actual business, complete with a public cash countdown and over 680 self-learned rules, all viewable live at firmulate.com/live. Decisions are made daily, recorded, and analyzed, providing transparent insights into AI’s practical capabilities.
The Deep Dive: Opus 4.8’s Profile
Among the four, Opus 4.8 was the most thorough participant, analyzing over 80 rules and performing deep diagnosis. Despite this, it left the deal on the table, showing a discipline slip—creating a warning that thoroughness alone doesn’t guarantee success in management tasks. Similar weaknesses appeared across all models, emphasizing that disciplined, comprehensive decision-making remains a challenge for AI.
What It Means for Your Business
For garden centers, outdoor retailers, or any outdoor living business, the takeaway is clear: AI can identify crises, resist manipulation, and even close deals. But it’s not just about chat quality—it’s about real performance under pressure, reading internal documentation, and staying disciplined. The league table shows that even newcomers like Kimi K3 can outperform established models when tested in real-world scenarios.
The Fairness Note
It’s worth noting that K3 ran without an effort parameter (the default API setting), while the others ran at xhigh. This demonstrates that different configurations can influence AI’s performance in management tasks.
Why Picking the Right Model Matters
Choosing an AI model isn’t just about flashy demos or chat prowess. As this experiment proves, the real measure is whether it can finish what it starts, read key internal documents, and stay honest—even in the face of manipulation. The league’s open competition makes it clear: the field is wide open, and testing models in your own business conditions is now more important than ever.
To see these models in action or to run your own business wargame, visit firmulate.com. Prepare your AI workforce before you hire it—and ensure it can deliver real results, not just talk.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
