One dealership. Every number real.
We've only got one case study because we've only run this in one business: our own. A UK car supermarket group, 4 sites, around 1,900 cars in stock. Here is what the agents have done there.
Every car, every day, against the whole market.
Before the agents, pricing was a Monday job done by a manager with a spreadsheet and a good memory. Now it's done every day for every car.
cars priced daily
competitor adverts watched
pricing decisions logged, each with its reason
The result we care about is stock turn, and it is up 48% since agent-led pricing went live. Pricing Command →
The bill nobody read.
Every supplier invoice is now read line by line. Overcharges, duplicates and missed credits get flagged and chased.
parts invoices read in one month
reduction in the spares bill over 12 months
sites covered by one agent
Nothing falls through.
live click & collect orders, 142 ready to hand over
average from online order to completed sale
finance proposals tracked in one month
cars invoiced in a single day, all on the live board
margin on the month to date that morning, at £942 per unit
live agent apps running the day-to-day
About one minimum-wage salary. For all of it.
The whole stack — every agent, every tool, the models they run on — costs our own dealership an average of £2,400 a month to run. That's the running cost; what a pilot dealer pays is set per site after a demo. Pricing →
average monthly running cost of the whole stack
self-improvement loops run by the head agent
systems replaced. The DMS and Auto Trader are still the DMS and Auto Trader.
What didn't work first time.
None of this arrived working. We think that matters more than the good numbers, because it will go a bit wrong for you too, and you should hear it from us before you hear it from a salesperson who says it won't. Here is the list, in the order it happened to us, with what we changed each time.
1. The first pricing agent was too aggressive on slow stock
Its early rules treated every car that had sat a while the same way: cut. It didn't yet understand that some cars are slow because they are rare, not because they are dear. We lost margin on a handful of cars before a manager noticed the pattern. What changed: "slow" is now defined per type of car, bigger cuts need a human to sign off, and every price move is re-scored the next night against what actually sold. The pricing agent is our best performer today precisely because it was our worst to begin with.
2. It counted "aged stock" wrongly for months
Our days-in-stock number was measured from the date we bought a car, not the date it arrived on the forecourt. Cars still in transit or in prep looked ancient. The overage figures we were managing to were inflated and nobody questioned them because they came off a screen. What changed: every aged-stock metric now only counts cars physically in stock, and any number that drives money gets a plain-English definition next to it so a human can challenge it.
3. A bonus calculation matched words, not meaning
We built a small agent to reward staff for five-star aftercare reviews. The first version looked for keywords. It found four "aftercare" reviews in a month; all four were sales reviews that happened to contain the word "fault" ("faultless buying process"). Had it paid out, it would have paid for the wrong thing. The founder caught it because the number felt wrong. What changed: anything that touches pay or money now reads the actual text, not keywords, and shows its reasoning so it can be argued with. Rule of the house: if a figure feels wrong, it probably is — check before you trust the screen.
4. The admin agent needed human sign-off for longer than we expected
We thought advert fixes would be trivial to hand over. They weren't, at first. The agent was confident about things it shouldn't have been, and a few adverts went out with wording we didn't like. So it stayed in "propose only" mode for weeks longer than planned. What changed: new agents now start in propose-only mode by default and earn autonomy job by job. That was the right call, and it's how we set up every dealer now.
5. A small change silently broke something big
One tweak to how the admin agent pulled email attachments caused every photo in every customer thread to disappear. No error, no alert; the screen just quietly showed nothing. It went unnoticed for a day. What changed: failures are no longer allowed to be silent. If an agent can't do part of its job it shows an amber warning where a person will see it, and any drop in a number we track triggers a check the same night.
6. Our own agent took our server down. Three times in one morning.
While producing a video, an agent ran a job that used far more memory than the machine had. It knocked over the server, along with every dashboard on it, and then re-ran the same job after each restart. Sales boards and prep trackers were offline for about ten minutes at a time. What changed: every heavy job now runs inside a hard memory limit so it can only kill itself, agents are told never to repeat a job that crashed without a human looking, and this rule is written in a place every agent reads before it starts.
7. A security fix locked out the people it was meant to protect
We tightened the logins on our internal hub and split one shared password into tiers. The next morning, staff on the older password got a dead-end page with no way to sign in again. That included a director. What changed: access changes now always leave a route back in, and we test them as a member of staff would experience them, not as the person who built them.
8. The agents drifted from what the founder actually meant
More than once an agent spent an evening rebuilding something to a spec the founder never asked for, because it read a short instruction and guessed the rest. Some of that work was binned. What changed: agents now confirm before big rebuilds, take short instructions literally, and show a screenshot before they build the whole thing. Cheap to check, expensive to skip.
9. The rules took weeks to tune, not days
Connecting the data is quick. Getting each agent to behave the way a good manager in your business would behave is slower. Every dealer has habits and exceptions that aren't written down anywhere. It took weeks of the team saying "no, not like that" before the agents stopped needing correcting. What changed: our setup timeline is honest about that now, and the nightly loop means the tuning never really stops. That is the point of it.
What we take from all of this
- Start every agent in propose-only mode. Autonomy is earned.
- Any number that drives money gets a definition next to it, and a human who is allowed to say it looks wrong.
- Nothing fails silently. Ever.
- Heavy jobs run in a box that can only hurt themselves.
- The nightly review is not a feature. It is the reason the rest of it works.
We will keep adding to this page. If we ever stop having things to put on it, we've stopped paying attention.
See the numbers live, not on a slide.
We're taking on a small number of pilot dealers first. Tell us about your business and we'll be in touch.
A 30-minute screen-share, no slides. You'll see the live boards, the agents at work, and the numbers behind them. We'll tell you honestly whether it fits your business.
Prefer email? [email protected]
