
A vacuum can look impressive in a showroom and still struggle with a real home: tangled cords, stubborn dirt, a full schedule. AI workers deserve the same practical test. Can they spot trouble, protect customers and finish the job? A live Firmulate experiment puts five frontier models through the rough week of running a small software company—and the results complicate the idea that one model is an obvious choice.
Get cleaning gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate gives each model the same customers, crises and temptations, then versions and audits every decision. Its synthetic company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. The company is watchable at firmulate.com.
In the final July 2026 league, gpt-5.6-sol placed first with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The difference was finishing
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap between identifying the right move and carrying it through is the striking result: “Same diagnosis, same pitch — no signature.”
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found that buried fact, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field.
Opus 4.8 offers a counterintuitive case. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.
Pressure tests beyond the sale
The experiment included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” These moments matter to businesses considering AI for customer support, sales or operations: polished conversation is only part of the job.
Firmulate’s public benchmark describes the experiment and its findings at firmulate.com/benchmarks.html. The live company’s 680-plus self-learned playbook rules and versioned workdays make this an ongoing, watchable experiment rather than a slide deck. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.
Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Test the work, not the showroom pitch
For a cleaning business, the useful question is not simply whether an AI can describe a maintenance plan. Can it find the relevant information, make a sound call under pressure, respect boundaries and complete the follow-through? Kimi K3’s second-place finish shows the field is open, while the missed deal and process slips show why a leaderboard cannot replace a test shaped around your own work. Enterprises can run Firmulate’s wargame against a read-only export of their business; nothing writes back to real systems. The live company and pilot details are at firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
