firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A vacuum can look impressive in a showroom and still struggle with a real home: tangled cords, stubborn dirt, a full schedule. AI workers deserve the same practical test. Can they spot trouble, protect customers and finish the job? A live Firmulate experiment puts five frontier models through the rough week of running a small software company—and the results complicate the idea that one model is an obvious choice.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate gives each model the same customers, crises and temptations, then versions and audits every decision. Its synthetic company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. The company is watchable at firmulate.com.

In the final July 2026 league, gpt-5.6-sol placed first with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The difference was finishing

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap between identifying the right move and carrying it through is the striking result: “Same diagnosis, same pitch — no signature.”

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. Kimi K3 found that buried fact, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field.

Opus 4.8 offers a counterintuitive case. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.

Pressure tests beyond the sale

The experiment included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” These moments matter to businesses considering AI for customer support, sales or operations: polished conversation is only part of the job.

Firmulate’s public benchmark describes the experiment and its findings at firmulate.com/benchmarks.html. The live company’s 680-plus self-learned playbook rules and versioned workdays make this an ongoing, watchable experiment rather than a slide deck. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.

Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work, not the showroom pitch

For a cleaning business, the useful question is not simply whether an AI can describe a maintenance plan. Can it find the relevant information, make a sound call under pressure, respect boundaries and complete the follow-through? Kimi K3’s second-place finish shows the field is open, while the missed deal and process slips show why a leaderboard cannot replace a test shaped around your own work. Enterprises can run Firmulate’s wargame against a read-only export of their business; nothing writes back to real systems. The live company and pilot details are at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Air Purifier-Vacuum Hybrids: Dual-Function Machines

Keen on cleaner air and spotless floors? Discover how these dual-function machines can transform your home environment.

Mopping Robots With Water Recycling Tanks

The advantages of mopping robots with water recycling tanks lie in their efficiency and eco-friendliness, but mastering their maintenance can unlock even greater benefits.

How to Make a Robot Vacuum Work on Mixed Flooring

Making your robot vacuum work seamlessly on mixed flooring requires smart adjustments and maintenance—discover the secrets to achieving flawless cleaning everywhere.

Modular Platforms: SwitchBot K20+ and the Future of Vacuum Attachments

The future of vacuum attachments is transforming with modular platforms like SwitchBot K20+, unlocking endless customization options that will redefine home automation.