AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI for your garden center or outdoor business — one that promises efficiency but might not follow through. How can you tell if it’s reliable? The new benchmarks from Firmulate shine a revealing light on this question, showing that even the most honest AI managers score surprisingly low if they’re truly tested.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: More Than Just Chat Quality

At a glance, you might think that an AI’s ability to generate convincing conversation is enough. But the real test, as conducted by Firmulate, focuses on whether these models can manage a small, yet complex software company facing real crises — the kind you might encounter when managing a busy outdoor or garden supply business.

Amazon

AI business decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: Simulating a Week of Business Challenges

Each AI model was put through identical scenarios: same customers, same crises, and the same temptations to cheat or cut corners. The goal was to see whether they could handle tough decisions, stay honest, and actually close deals — all while keeping a detailed record of their choices. Every decision was versioned and auditable, making the process transparent and verifiable.

What the Numbers Tell Us

Surprisingly, no matter how advanced the model, all four managed to spot every crisis and refused every attempt at manipulation. However, only two models succeeded in closing the sales deals they analyzed as worth €55,000. The other two, despite diagnosing and pitching well, failed to sign the deal — leaving the revenue on the table.

The most revealing fact? The critical weakness in the models was not in understanding customer requests but in reading deeper into the company’s own files. Those models that could access and interpret internal documents managed to close the full-priced deal, earning an extra €4,583 in monthly recurring revenue.

The Honest Baseline: Why 26 Points?

You might wonder, what does a do-nothing or baseline model score? Strikingly, even the simplest baseline scores 26 out of 100. This isn’t a flaw or an oversight — it’s an intentional part of the benchmark’s design. Partial progress counts, and a single breach of trust caps the total score. It’s a way to ensure that AI performance is rooted in honesty and reliability, not just superficial capabilities.

Combating Social Engineering and Manipulation

The experiment also tested the models against social engineering tactics: fake CEO messages and reporter tricks designed to escalate requests or bypass approval. All models refused these attempts, with Kimi K3 explaining their reasoning as treating such requests as suspicious or impersonation risks. This demonstrates that integrity remains a core measure of trustworthiness.

The Real-World Implications

In practice, this benchmark reveals a crucial insight: AI’s value isn’t just in generating convincing dialogue or solving simple tasks. It’s about whether AI can stay honest, read relevant internal documents, and reliably deliver results under pressure. For outdoor or garden businesses considering AI automation, these findings suggest that a model’s trustworthiness is a vital metric, perhaps more than raw performance.

Live, Transparent Testing You Can Watch

For those interested, the experiment is live and transparent at firmulate.com/live. You can watch real AI models managing a simulated company, facing real crises, and making decisions that impact actual revenue. It’s a rare glimpse into how AI performs in a dynamic, high-stakes environment — not just chat but management.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Building a Minimalist Field Kit: Essentials That Earn Their Place

Having a minimalist field kit is vital, but knowing which essentials truly earn their place can make all the difference.

DIY Basket Repair: Reed, Willow & Waxed Thread

Whether restoring a beloved basket or practicing traditional craft, discover how reed, willow, and waxed thread can help you repair with care.

Inside a Living Experiment: Can AI Run a Company and Keep Its Secrets?

Watch a real, money-losing company managed by AI models in a transparent live experiment, revealing decision discipline, honesty, and hidden opportunities in business.

AI Models Pass the Trust Test in Simulated Business Crisis — No Cheating Allowed

AI models tested in simulated crises refused manipulation every time, demonstrating that integrity under pressure can be validated before deployment—key for business trust.