The most common state of an enterprise AI initiative is not success or failure. It is purgatory: a pilot that demos well, offends no one, proves nothing, and never ends. Six months in, nobody can say whether it worked, because nobody defined what working meant. The pilot is not dying. It is worse — it is undead, consuming budget and credibility while blocking the honest conversation a failed experiment would have forced.

I have been brought in behind enough of these to see the pattern clearly. Pilots do not stall for technical reasons. They stall for product-management reasons that were locked in before the first line of code.

The six pre-conditions of pilot purgatory

1. Nobody answered “how good is good enough.” Every stalled pilot shares this. The team can tell you the model, the vendor, the architecture — but not the accuracy threshold at which the business would act. Without a pre-committed quality bar, every result is interpretable, and interpretable results extend pilots forever. The bar is a business decision, not a data-science one: what error rate would you accept from a new hire doing this task?

2. There is no baseline. “The AI resolves tickets in four minutes” means nothing if nobody measured the human process first. Teams skip baselining because it is boring and slows the kickoff. Then they spend month five arguing about whether four minutes is good — an argument that costs more than the two weeks of measurement would have.

3. The pilot runs beside the workflow, not inside it. The sandbox pilot — a parallel tool, a separate tab, a spreadsheet export — measures whether the technology functions, which is not the question. The question is whether it survives the real workflow: interruptions, exceptions, handoffs, the CRM that hasn't been updated since the acquisition. Integration debt discovered in month five was visible in week one to anyone who walked the actual process.

4. No operator owns it. Pilots sponsored by innovation teams and delivered to no one die the day the innovation team's attention moves on. The owner must be the person whose team lives in the workflow — the one whose numbers improve if it works and who will be embarrassed if it is quietly abandoned.

5. Adoption was assumed, not designed. Frontline staff do not distrust AI because they fear technology. They distrust it because they own the error. If the tool is wrong and the customer is angry, the human who accepted the output takes the consequence. A pilot that does not design the trust ramp — visible confidence signals, easy correction, a fallback path that costs nothing to use — will report “low engagement” and conclude, wrongly, that the technology wasn't ready.

6. Success was defined as sentiment. “The team loves it” is the most dangerous sentence in enterprise AI. Sentiment is real but unbankable. Pilots need one primary metric that moves a business number someone already reports upward — handle time, first-pass yield, days outstanding — and the discipline to ignore applause that arrives without it.

A pilot without a kill criterion is not an experiment. It is a subscription — one you pay in credibility as well as budget.

The 90-day pattern that survives

The structure I run — refined across SaaS, healthcare, and automotive-retail engagements — compresses to four phases with hard gates:

Weeks 0–2: Baseline and definition of done

Measure the human process as it actually runs: cycle time, error rate, exception rate, cost per unit. Write the contract before building: the quality bar, the primary business metric, the kill criterion, the scale criterion. All four signed by the operating owner — not the innovation sponsor, not the vendor. If you cannot get signatures in two weeks, that is the pilot's result. You saved a quarter.

Weeks 3–6: The narrow wedge, in production

Pick the thinnest slice of the real workflow — one document type, one ticket category, one region — and run it live with a human checkpoint on every output. Narrow is the point: a wedge that handles 8% of volume inside the real workflow teaches more than a sandbox that handles 80% beside it. This is where integration debt surfaces while it is still cheap.

Weeks 7–10: Instrument quality, cost, and trust

Three curves, tracked weekly: quality against the bar, fully loaded cost per completed task, and human acceptance rate — how often the checkpoint accepts the output unchanged. The third curve is the one teams skip and the one that predicts scale. Quality that isn't trusted delivers nothing.

Weeks 11–12: The forced decision

The gate that defines the whole pattern: on the pre-committed date, against the pre-committed criteria, the pilot scales, kills, or — rarely, and only with a named reason — extends once. The decision is made in an operating review, on the numbers, by the owner who signed in week two. No new evidence is admitted that wasn't defined at the start; that rule exists because purgatory is built from mid-pilot goalpost moves.

Killing pilots is a capability

The organizations that get durable value from AI are not the ones whose pilots always succeed. They are the ones whose pilots always conclude. A clean kill in week twelve — with a baseline, a documented gap, and a written reason — is an asset: it prices the capability curve for your context, and it tells you exactly what has to become true before you try again. An undead pilot teaches nothing and salts the ground for the next attempt, because the second pitch for AI in a workflow is made to an audience that remembers the first.

Demos are cheap now. Decisions are the scarce resource. Structure your pilots so a decision is the guaranteed output, and the technology results will take care of themselves.