firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI know when to close the deal?

Imagine an AI helping run a busy coffee business: customers are changing their orders, a competitor is making moves, and someone claiming to be the CEO asks for a quick exception. A polished answer in a chat window is easy. Making the right call across a messy week of business is harder. Firmulate’s live experiment puts that difference on display.

A company under pressure

Firmulate put frontier AI models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Decisions were versioned and auditable, and the experiment is watchable at Firmulate.

The company is synthetic, but its money mechanics are real: 13 synthetic employees, €105,000 in monthly burn against €2,300 in monthly recurring revenue, and a public cash countdown. Its evolving playbook contains more than 680 self-learned rules, with every workday versioned. The point is not to pretend this is a coffee shop. It is to watch how AI handles the kinds of decisions any business may face before handing it a real role.

Seeing the crisis is not the same as acting

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand for that gap: “Same diagnosis, same pitch — no signature.” A model can describe the right move and still leave the opportunity on the table.

The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That detail makes the test feel less like a quiz with obvious clues and more like a real workplace: useful evidence may be present, but someone has to find and use it.

Trust was tested directly, too. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusing pressure matters, but the deal showed that caution alone does not complete the job.

The leaderboard has a caveat

The final Crucible League standings for July 2026 put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Those results come with a fairness note: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Opus 4.8 offers a useful reminder that effort and thoroughness do not guarantee a strong outcome. It was the most thorough participant, with 80 additional learned rules and the deepest analyses, but came last. It left the deal unsigned and tried writing into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.

For readers curious about how managers make decisions, Firmulate also has a “guess the model” quiz built from 242 real, unedited management decisions. It is available at firmulate.com.

From watching to trying it at work

A live company can make these patterns easier to see. A pilot asks a more practical question: how would an AI respond to crises using information from your business? Enterprises can run the wargame against a read-only export, test scenarios against their own company, and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Take the question into your own business

Watching an AI handle a simulated company can reveal a gap between sound analysis and follow-through. A pilot can show what that gap looks like against your own business data, without writing to live systems. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What to Expect at the 2025 World Barista Championship

Lurking behind the scenes at the 2025 World Barista Championship are groundbreaking techniques and trends that will redefine coffee’s future—keep reading to discover more.

Whataburger Surges In Global Coverage

Whataburger experiences a surge in international media mentions, with GDELT reporting a tenfold increase in coverage this week.

Mountain Dew Will Release a 5-Cent Limited-Edition Soda — How to Get It

Mountain Dew will offer a limited-edition soda for 5 cents. Here’s how customers can get it and what to expect from this promotion.

Texas Egg Recall Over Salmonella

Texas officials recall eggs over salmonella risk, affecting multiple brands. No confirmed reports of illnesses so far. Details on affected products and next steps inside.