
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
What a simple, do-nothing AI can teach your garage about trust and performance
In the fast-evolving world of automotive services, trusting AI to handle customer interactions, scheduling, or inventory might seem straightforward. But a recent benchmark experiment by Firmulate reveals a surprising truth: even the most passive AI baseline scores 26 out of 100. That’s not a typo. It highlights how cautious, honest performance is measured, and why it matters for your business.
As an affiliate, we earn on qualifying purchases.
The Science Behind the Score
Imagine running your garage’s management tasks through an AI that does absolutely nothing — no reading customer emails, no ordering parts, no scheduling repairs. Yet, this baseline AI still scores 26 points. How? Because in the experiment, partial progress counts. For example, recognizing a crisis or identifying a problem, even if no action is taken, boosts the score. Conversely, a single breach of trust — such as attempting to manipulate a customer or cut corners — caps the total score at 26 and prevents further improvement.
This approach is deliberate. The benchmark is designed to measure honesty, discipline, and reliability — the core qualities that make AI trustworthy in real-world operations. It’s not just about how well an AI can generate natural language but whether it can stay honest, follow protocols, and prioritize customer interests even under pressure.
Real-World Insights from the Experiment
Firmulate’s live experiment involved four frontier AI models, each managing a simulated small software company facing its worst week. The models faced identical crises, customer demands, and ethical temptations. Every decision was versioned and auditable, mirroring real operational conditions.
Remarkably, all four models identified every crisis and refused manipulation attempts, including complex social engineering tactics like fake CEO messages and journalists’ tricks. For example, when fake requests escalated in multiple stages, all models refused to approve anything suspect, citing reasons like ‘possible impersonation.’
However, only two of the four models managed to close the deal worth €55,000 — the same diagnosis and pitch, but only these two signed the contract. The others left the deal on the table, despite recognizing the opportunity, due to discipline slips such as writing requests into locked departments instead of escalating them.
What Really Wins the Deal? Reading the Files
The critical weakness identified was in how models accessed and used internal company documents. The models that read two document references deep into the company’s own files won the deal at full price, worth more than €4,583 MRR. This underscores a vital lesson: trustworthiness isn’t just about surface-level performance but about thoroughness and access to critical information.
Social Engineering and Ethical Vigilance
In scenarios designed to trick AI into breaching ethics — like pretending to be the CEO or requesting background approvals — all models refused. Kimi K3 explained its refusal by treating such requests as ‘suspected approval-bypass or possible impersonation.’ This consistent stance demonstrates that these models can be programmed to prioritize integrity even in complex social manipulations.
The Practical Reality of Managing a Garage
The live experiment runs at firmulate.com/live showcase a real money environment with 13 synthetic employees, daily decision-making, and a monthly cash burn of €105,000 against a revenue of €2,300. This is not just an academic exercise; it’s a simulation of real, high-stakes management. The models are supported by over 680 self-learned rules, and every decision is versioned for transparency.
For your garage, this means that deploying AI isn’t just about automation or customer engagement. It’s about ensuring your AI can finish what it starts, read and understand critical documents, and stay honest under pressure. A model that can slip into process slips or ignore internal files isn’t just less effective — it could cost you deals or damage trust.
Understanding the Scores and Their Implications
In the recent leaderboard, the top model scored 95 out of 100, having found the hidden fact and secured the full deal. The second scored 93, the third 88, and the fourth 77, with each showing some process slips. The baseline, by contrast, scores 26, illustrating that even a do-nothing approach garners partial credit for recognizing issues, but full trustworthiness requires more.
Importantly, the benchmark emphasizes that trust isn’t optional. A single breach caps the score, making honesty and discipline non-negotiable. For garage owners, this underscores the importance of choosing AI solutions that are transparent, disciplined, and capable of thorough information analysis.

Key Takeaway: Trust and Discipline Are Non-Negotiable in AI
In deploying AI for your garage, the lesson is clear: it’s not enough for AI to generate decent responses. It must read, understand, and uphold honesty under pressure. The benchmark’s revealing score of 26 for a do-nothing baseline reminds us that trustworthiness is the foundation of effective AI — and often, the hardest quality to teach but the easiest to lose. Investing in trustworthy AI means prioritizing systems that can finish what they start, read critical files, and refuse manipulation, ensuring your business remains honest and reliable in every transaction.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
