On a Friday afternoon, the reordering agent did exactly what it was built to do. A mid-market manufacturer had wired it up over a quiet few months. When stock on a part dropped below its reorder point, the agent drafted a purchase order, checked it against the preferred-supplier list and the current price file, and, for six weeks, handed it to a buyer to approve. It was right almost every time. So someone asked the reasonable question: if it’s right almost every time, why is a person still clicking approve?

They took the person out of the loop on a Friday.

Over the weekend, an upstream price file came in malformed, and the agent read a decimal in the wrong place. Every order it placed was individually plausible. Collectively, by Monday, they were a mess nobody had approved — because the entire point of the change was that nobody had to.

Nothing about the model was wrong. The model did its job perfectly. What failed was the rollout.

The Word That Changes Everything

Every AI tool before this one answered. You asked, it replied, and a human decided what to do with the reply. The human was the gap between the machine and the consequence.

Agentic AI closes that gap. It doesn’t answer — it acts. It calls tools, writes records, places orders, moves work forward while you’re not looking. In practice, that capability arrives through something like MCP — a standard way to hand an AI a set of tools and let it use them. The protocol is the easy part. The discipline is deciding which tools, when, and with what-happens-if-it’s-wrong already built.

That one word, acts, is why you don’t roll out agentic AI by flipping a switch. You roll it out on a leash you lengthen as it earns your trust. There are three lengths.

Rung one: Observed

The agent gets read-only tools. It can see the stock levels, read the price file, look at the open orders — and it drafts what it would do. A human does every actual write.

This rung feels pointless to skip past. It isn’t. It’s where you learn what the agent gets wrong before “wrong” can touch anything. You are not evaluating whether it’s clever. You are building a record of how often its proposal matched what the human actually did.

You’ve earned the next rung when that match rate is boring — high, stable, and you understand the misses. Not “it seems good.” A number you’d defend.

Rung two: Supervised

Now the agent gets the tools that write — but every action it takes routes through a human approval before it commits. The agent drafts the PO and places it, pending one click.

The failure here is subtle and worth naming: approval fatigue. When the agent is right ninety-odd times in a row, the hundredth approval gets rubber-stamped, and rubber-stamping is just unattended mode with extra steps and a scapegoat. If your reviewers are clicking approve without reading, you don’t have supervision. You have theater. Watch for it.

You’ve earned the next rung when the approvals have become genuinely uneventful and you can answer the next question cold.

Rung three: Unattended

The agent acts on its own, inside a fence. No per-action approval. This is the rung everyone wants on day one and should reach last.

Before you unclip the leash, answer the one question the Friday-afternoon crowd skipped.

The Question Nobody Scopes: BLAST RADIUS

Not “is the agent accurate?” You already know that from rung one. The question is: what is the worst thing one bad unattended run can do with the tools it’s holding?

Scope the tools, not just the prompt. An agent told to “only reorder normal quantities” but handed an unlimited PO tool is fenced by a sentence. An agent whose ordering tool physically cannot exceed a dollar ceiling, a quantity cap, or the approved supplier list is fenced by a wall. Sentences are suggestions. Walls are the rollout. Decide the blast radius, then build the fence in the tools before the agent ever runs unattended.

The Kill Switch Is A Feature, Not A Reaction

The Friday story ends better if one thing was true: someone could shut the agent off — instantly, without the agent’s cooperation — and the workflow fell back to the way it worked the week before.

That is a feature you design on day one, not a thing you improvise at 6 a.m. Monday. Revoking an agent’s tool access should be one action, it should not depend on the agent behaving, and the fallback should be the boring old process, still warm. If turning it off is hard, you didn’t roll out an agent. You installed a dependency.

The Posture

The shops rolling out agentic AI well are not more advanced than the ones getting burned. Often they’re using the same tools. They’re just slower to hand over the keys — deliberately, on a schedule, against evidence they can point to.

Earn the leash. Then lengthen it. The technology will be ready to act long before your process is ready to let it, and the gap between those two dates is the whole job.

Plant Floor to Cloud goes out every Tuesday.

Rolling out something that acts on its own — or cleaning up after one that did? Reply, I read every one.

Reply

Avatar

or to participate

Keep Reading