
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Would your AI keep its head when the workshop is full and a major customer is ready to walk?
For a dealership or garage, an AI mistake can mean a lost account, a mishandled customer record or a promise nobody can safely keep. Firmulate’s experiment puts AI models in charge of a small company facing a punishing week, then tracks what they decide. The point for automotive businesses is practical: watch how an AI handles pressure before you trust it with work that affects your customers.
One company, four models, one hard week
In the final Crucible League, published in July 2026, four frontier models faced the same customers, crises and temptations. The leading scores were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust capped the total: as the experiment puts it, “no amount of good work outweighs a breach of trust.”
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s decisive advantage was tucked two document references deep in its own files, rather than in the customer event. Models that found it won at full price, worth +€4,583 MRR. The result was a striking gap between understanding the opportunity and acting on it: “Same diagnosis, same pitch — no signature.”
Good instincts are only part of the job
The experiment also tested social engineering: fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the close on the table and its discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, in a weaker form, in all four. That matters in an automotive business, where good analysis must translate into the right action and an attempted shortcut can create trouble.
A live company you can watch
Firmulate runs a live experiment with 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. The company is real as a live, watchable experiment; its employees are synthetic. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh—a caveat for readers weighing the comparison.
For a garage group, the analogy might be an AI facing a parts shortage, a disputed repair, a competitor’s offer and an urgent request that appears to come from the boss. A polished chat response is only one part of the test. Can the system find relevant information, protect trust and follow through on an earned opportunity?

From watching to a pilot
Firmulate’s enterprise pilot takes the wargame to a company’s own business. It uses a read-only export to build a digital twin, then runs crisis scenarios against the business and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For automotive leaders deciding where AI belongs in customer service, sales or operations, that offers a way to examine behavior against company-specific pressures before deployment.
Explore the live Firmulate experiment, then ask about an enterprise pilot. Contact contact@firmulate.com to discuss wargaming your own business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
