AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A greenhouse can look healthy right up to the moment a pump fails, a shipment is delayed, or a customer calls with a problem. In a busy growing season, the test is not whether a plan looks good on paper. It is whether the people—and tools—you rely on notice trouble, protect trust, and follow through when several things go wrong at once.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is the practical question behind Firmulate, an experiment that puts AI models in charge of a small software company during a simulated worst week. The setting is not a garden business, but the challenge will feel familiar to anyone who depends on coordinated decisions under pressure: the models face the same customers, crises, and temptations, then their choices are recorded for review.

A shared week, with real consequences inside the experiment

Firmulate’s final Crucible League, completed in July 2026, compared five participants. Each frontier model ran the same company through the same difficult week. Every decision was versioned and auditable, so observers could follow what happened rather than judge a polished demonstration after the fact.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rules count partial progress, but a single breach of trust caps the total. The principle is direct: “no amount of good work outweighs a breach of trust.”

Across the exercise, all models spotted every crisis and refused every manipulation attempt. That sounds reassuring, but recognition was only part of the job. Only two models signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” The difference between identifying an opportunity and completing the work is easy to miss in a conversation demo. Here it showed up in the decisions.

The important clue was already in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not spelled out in the customer event that triggered the sales opportunity. Models that read the relevant file won the deal at full price, worth +€4,583 MRR.

For a grower, the parallel might be a customer’s delivery history, a note about a supplier, or a detail in an operating procedure that changes what to do next. The lesson from Firmulate is not that every business has the same hidden clue. It is that an agent’s answer to the visible problem may depend on whether it consults the useful information the company already holds—and whether it completes the action that follows.

The experiment also tested attempts to exploit authority and urgency. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a promising sign for teams concerned about agents being pressured into disclosing information or bypassing normal approvals.

Thoroughness did not guarantee the best result

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and its discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models.

That mismatch matters for businesses weighing where to use AI. A careful explanation is not the same as a completed task, and recognizing a boundary is not the same as responding appropriately when an action is blocked. A greenhouse operator would still want to know whether a system can move from noticing a temperature alert to following the right escalation path—not merely describe the alert convincingly.

There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published result should be read with that condition in view.

Watch the company, then test your own

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and a versioned record for every workday. The company is synthetic, while the financial pressure and decision trail make the experiment watchable as it unfolds. The public Firmulate site presents the live experiment alongside its league and other ways to explore the results.

There is also a “guess the model” quiz built from 242 real, unedited management decisions. It invites readers to judge the choices themselves before seeing which model made them. Together, the live company and decision record shift the discussion from whether AI sounds capable to what it actually does across a chain of events.

Firmulate’s next step is aimed at enterprises: run the same kind of wargame against a read-only export of the company’s own business. The exercise can put company-specific crisis scenarios against existing playbooks and produce a board report with model rankings and weaknesses to examine. The export is read-only; nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to acting

A public experiment can reveal patterns, but each business has its own customers, procedures, and pressure points. A pilot lets an enterprise examine how models handle scenarios grounded in its own information while keeping the exercise separate from live systems. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Beat the Heat: Top Tips for Using the Ninja Blast Max Portable Blender This Summer

Get the most out of your Ninja Blast Max Portable Blender with our summer hacks for fresh smoothies, frozen drinks, and on-the-go refreshment.

Summer Gatherings Made Easy with the Ninja Creami XL Ice Cream Maker

Host summer parties effortlessly with the Ninja Creami XL. Make creamy ice creams, sorbets, and more for family and friends with ease.

DEWALT 20V MAX Combo Kit vs DEWALT ATOMIC Drill: Full Comparison

Compare the DEWALT 20V MAX Combo Kit with the DEWALT ATOMIC Drill to find the best cordless drill for your needs. Detailed specs, pros, cons, and recommendations.

DEWALT vs Ryobi: Which Power Tools Wins in 2026?

Compare DEWALT’s cordless drill with Ryobi’s lineup to find the best power tool for your needs in 2026. Discover which brand offers better performance and value.