Two months after go-live, the AI dashboard is a wall of green. Accuracy: 94%. Queries handled: up and to the right. Average response time: under a second. Everyone who built it is pleased.

And the business case is quietly falling apart, because not one of those numbers measures whether the deployment is doing its job.

That’s the trap of demo metrics. They measure the model — how clever it is, how fast, how much it gets used. They were chosen to impress a room, and they keep impressing long after the deployment has stopped delivering. Deployment metrics measure the outcome: whether the work is getting done faster, cheaper, and correctly enough to trust. The two sets rarely move together, and when a project is in trouble, the demo metrics are the last thing to show it.

Issue #5 covered which AI projects earn the word “deployment” before you fund them. This is the other half: once one is live, the four metric swaps that tell you the truth.

Swap 1: Model accuracy → Exception rate, and its trend

“94% accurate” is a demo metric. Accurate against what — a benchmark, a test set, last quarter’s data? And what happens to the other 6%?

The deployment metric is the exception rate: how often a human has to step in and override or correct the output. Track it as a trend, not a snapshot. A flat exception rate means the model is holding. A creeping one means drift — the world changed, and the model didn’t — and it will keep looking “94% accurate” on the old benchmark the whole time it’s going wrong on live work.

Swap 2: Usage → Task completion

“Engagement is up. 3,000 queries last month.” Usage is the most seductive vanity metric there is, because it always goes up when a tool is new and always feels like success.

The deployment metric is task completion: of the work this thing was bought to finish, how much actually got finished — start to done — without a human redoing it. High usage with low completion isn’t adoption. It’s people poking at a tool that doesn’t quite work and then doing the job themselves anyway.

Swap 3: Response speed → End-to-end cycle time

“Sub-second responses.” Impressive, and almost never the bottleneck. If the model answers in 800 milliseconds and then a person spends nine minutes checking the answer, you did not build a fast process.

The deployment metric is end-to-end cycle time, human review included — from the moment work enters the workflow to the moment it’s done and trusted. That’s the number the business felt before the AI and the only one it will feel after. Optimizing model speed while ignoring review time is optimizing the part that was never slow.

Swap 4: “Time saved” → Realized cost per unit, and who measured it

“Saves about 20 hours a week.” Saved by whom, measured how, and did that time turn into anything the business can bank? A hallway estimate is a demo metric wearing a number’s clothing.

The deployment metric is realized cost per unit of work — per case, per order, per document — measured before and after, by someone who isn’t the person who pitched the project. If nobody owns that measurement, you don’t have a payback; you have a hope. This is the one that moves a budget decision, which is exactly why it’s the one that goes unmeasured.

The one-line test

For any metric on the dashboard, ask: could this number, moving the wrong way, actually change a funding decision? If yes, it’s a deployment metric — keep it. If it could go red while the budget stays comfortable, it’s a demo metric. Demo metrics aren’t useless; they’re diagnostics for the engineers. They just don’t belong on the scorecard you defend to the people who paid for the thing.

The cost of watching the wrong numbers

Optimize demo metrics and you get a deployment that is provably excellent at everything except paying for itself. It sails through every review on a green dashboard while the exception rate climbs, the humans quietly route around it, and the cost-per-unit never actually drops.

Then one of two things happens. Either it limps for a year as an unkillable line item nobody will defend — or it fails somewhere real, loudly, and the post-mortem asks why the dashboard never warned anyone. It didn’t warn anyone because it was never wired to the outcome. It was wired to the demo.

The shops building durable AI capability aren’t measuring harder. They’re measuring the four things that move a budget, and letting the model’s vanity stats stay where they belong — in the engineering channel, not the boardroom.

Plant Floor to Cloud goes out every Tuesday.

Got an AI deployment whose dashboard lied to you — or one whose numbers actually held? Reply, I read every one.

Reply

Avatar

or to participate

Keep Reading