Two AI projects. Comparable budgets, comparable teams, started in the same quarter. A year later, one is quietly running in production and paying for itself. The other is a slide in a steering-committee deck that nobody opens anymore.
The difference is almost never the model. The model was the easy part. The difference is everything around the model — and it is visible, if you know where to look, before you spend a dollar.
The one that earns money
It solves a problem you can name in dollars and hours. Not "improves efficiency." A specific person was spending a specific number of hours a week on a specific task, and now they spend a fraction of that and review the exceptions. The savings were nameable at the proposal stage, which is why the project got funded by an operating budget instead of an innovation budget.
It has a boundary. The AI does one thing. The system around it constrains what goes in, and the output is checkable against something — a rule, a range, a prior value. When it is wrong, you can tell, fast.
It keeps a human on the consequential calls. Anything that touches a customer, a partner, or money has a person in the loop, not because the AI cannot draft it, but because the cost of a confident wrong answer is higher than the cost of the review.
And it has a way to fail safely. If the model degrades or the vendor has an outage, the workflow falls back to the way it worked before. Nobody finds out from an angry customer.
The one that doesn't
It was built to impress. The selling point was breadth — "it handles anything," "it works on unstructured input," "it'll learn." Breadth is the tell. Breadth means no boundary, and no boundary means no way to know when it's wrong.
Its value was described in adjectives, not numbers. When you ask what it saves, the answer is a feeling. That's because it was scoped to demo well, not to pay back, and demos reward range while operations reward reliability.
The human reviewer got removed in the name of efficiency — usually right after the demo, when someone asked why a person was still in the loop if the AI was so good. And there is no fallback, because nobody planned for the AI to be wrong; being wrong wasn't in the demo.
How to tell them apart before you fund it
Four questions, asked of the person proposing the spend, not the vendor selling it:
What does this save, in dollars or hours, and who measured it? If the answer is an adjective, stop.
What's the boundary — what is the one thing it does, and what's explicitly out of scope? No boundary, no deployment.
What happens on the day it's wrong? If there's no answer, there's no plan.
What's the fallback if it goes down? "It won't" is not a fallback.
A proposal that survives those four questions is a deployment. One that doesn't is a science project — and science projects are fine. They belong in an R&D budget, with an R&D timeline and R&D risk tolerance. The failure isn't running experiments; it's running experiments in an operations budget and calling them deployments, then acting surprised when they don't behave like deployments.
The cost of getting it wrong
The wrong AI project rarely fails loudly. It limps. It consumes a maintenance budget, a vendor relationship, and a slice of your team's credibility every time someone asks "didn't we buy something for this?" The real cost isn't the sunk build — it's the next three good ideas that don't get funded because the last AI project soured the board on the category.
The shops building durable AI capability aren't smarter about models.
They're stricter about which projects earn the word "deployment."
Plant Floor to Cloud goes out every Tuesday.
Got an AI project that earned its keep — or one that didn't? Reply, I read every one.

