Issue #10 was an admission: we shipped an AI chatbot and turned it off after eight weeks, because our buyers didn't want a conversation. Issue #11 asked who has the standing to make that call.
This one is different, and it's the one I still think about. Because this time the system worked.
Last year an early-warning capability on the quality side of the business was built. It read incoming customer sentiment, worked through the free-form text where the real complaints live, and watched patterns across resolution steps — hunting for the signature of a part starting to fail in the field before rising claim volume would make it obvious.
It existed for a reason. The business had already lived through a significant quality event. That event was why there was attention, and why nobody had to be sold on the idea of seeing the next one coming sooner.
So we tested it the way you're supposed to. We took two years of historical data, including that event (which the model had never seen), and asked whether it would have caught it.
It did.
And the solution still never reached production. A little over three months in proof of concept, and it died on the vine.
What Did We Get Wrong?
We thought the back-test was the finish line. We thought if we could show it caught a real event, the argument was over.
The argument was just starting, and the questions that ended it were not the ones we'd prepared for:
How did it arrive at that? What was the scoring method? Was that a lucky guess, or actual signal? And who, exactly, is going to explain that to someone who asks?
We had answers to none of them in a form a business team could use.
Here's the part that took me longest to accept: the back-test validated the answer. It did nothing for the reasoning. Those are different things, and I had quietly assumed the first would stand in for the second.
It doesn't. A correct result produced by a method you can't articulate isn't evidence — it's a narrative. And we had exactly one of them. One event, caught once. Statistically that is very close to nothing. If the model had found it by reading genuine signal in complaint language, we'd have seen the same outcome. If it had found it by latching onto some spurious correlation that happened to line up, we'd have seen the same outcome. From outside the black box, a working model and a lucky one are indistinguishable.
The quality and field teams understood this before I did. They weren't being obstructionist and they weren't technophobic — they validate findings manually all the time and did so throughout. Verification was never the gap. They were pointing at something I'd waved past: we cannot tell the difference between this system being right and this system being fortunate, and we are the people who will be asked.
The Part I'd underweighted: This Is a Brand Problem
The argument that finally landed for me wasn't technical at all.
If your brand promise is reliability — if the entire proposition is that customers can trust what you put in the field — then acting on a finding you cannot explain means spending trust you cannot account for. Pull inventory, open an investigation, contact the field on the strength of a score nobody can reconstruct, and you'd better be right. Be wrong once, publicly, on unexplainable grounds, and you have damaged the exact asset the company is built on. Worse than the original problem.
No employee wants to be the one who did that. So the finding gets acknowledged and not acted upon, which is the worst available outcome — the organization is now generating consequential information nobody will move on.
That reframed explainability for me permanently. It is not a technical nicety or a compliance checkbox. Explainability is a brand control. It's the mechanism by which a company gets to stand behind a decision, and a business whose promise is trustworthiness cannot make decisions it can't stand behind.
The Framework
Four things I'd now require before any inferential system goes to a team that has to justify its decisions:
An explanation a human can deliver out loud, in ninety seconds, without you in the room. Not feature-importance charts. A sentence: “it flagged this because complaint language around this part shifted in a way that historically precedes failures.” If that sentence doesn't exist, the system hasn't shipped — it's still a demo.
A falsification standard, defined before the back-test. What would this model have to miss for us to distrust it? Decide that in advance, or a successful back-test will simply confirm what you already wanted to believe.
A sample size agreed in advance that beats luck. One catch is not a trend. Name the bar before you run the test, because afterward everyone negotiates it.
A named owner of the “how does it work” answer — someone other than the person who built it, who can hold the room when challenged.
The Cost
The model was the cheap part. That's what's worth sitting with.
We spent our effort where the interesting technical problem lived — scoring, correlation, text analysis — and almost nothing on the explanation layer, because we'd assumed accuracy would carry it. Three months of a capability that appeared to work, ended by a question we never budgeted a single hour to answer.
The loss isn't the build. It's that something which may well have been genuinely valuable is shelved, and the next proposal in that space inherits the sentence we tried that, it didn't go anywhere. That sentence will not include the word explainability. The model takes the blame for a communication failure, and the next attempt is harder to fund than this one was.
Three Tuesdays later, the same lesson in different clothes. #10: the deployment that earns money is the one that matches how the user actually behaves. #11: somebody has to own the decision to stop. This one:
Accuracy buys you the right to be considered. Explainability is what buys adoption — and if your brand runs on being believed, you cannot act on what you cannot explain.
Plant Floor to Cloud goes out every Tuesday.
Have you had a model pass every test you designed and still fail the one you didn't think to run? I'd like to hear it. Reply, I read every one.

