The pilot succeeds, everyone is pleased, and nothing scales. This happens often enough to be a design problem rather than bad luck.
Pilots stall most often because they were designed to prove the technology works rather than to answer the questions that gate a production decision: who owns it, what it costs at scale, how it degrades, and whose workflow absorbs the change. A pilot that only tests feasibility will succeed and still leave the deployment decision unmade, because none of the blocking questions were in scope.
A pilot is scoped. A motivated team runs it in one department with vendor support close at hand. It works. The report is positive. Then the conversation about scaling begins, and it surfaces questions nobody was asked to answer: who funds it beyond the pilot, who maintains it, what happens when the champion moves roles, whether the workflow holds in a department that did not volunteer.
The pilot did not fail. It answered a question that was never the blocker.
In our experience the distinguishing feature is unglamorous: the successful ones established a baseline before they started, and picked a workflow where someone already owned the outcome. When the pilot ended they could say what changed against a number that existed beforehand, and there was a named person whose job got easier or harder.
The stalled ones usually cannot answer what the process cost before the tool arrived, which means the benefit case is an estimate arguing with a budget.
We would go further than most: a large share of healthcare AI pilots should not be run at all. If the organisation cannot articulate what decision the pilot will settle, and what result would cause it to walk away, the pilot is a procurement ritual with a technical costume.
The most useful discipline we give clients is to write the stopping condition first. Before the pilot starts, state the result that would make you decline to proceed. Teams find this uncomfortable, which is the point. A pilot with no failing grade is not an experiment.
The second discipline: pick something boring. The pilots that scale tend to address a workflow nobody enjoys, with a measurable baseline and a clear owner. Ambitious clinical pilots generate more excitement and far fewer production deployments.