Baselines are cheap before deployment and impossible afterwards. This is the least glamorous and highest-return hour of any AI programme.
Before deployment, capture four things: the time the current process actually takes measured from system logs rather than estimated, the current error or rework rate, a validated measure of staff experience for the people whose work changes, and the volume the process handles. Without a pre-deployment baseline, any later benefit claim is an estimate arguing with a budget, and the organisation cannot tell a working deployment from a popular one.
Baselining is boring, it delays the exciting part, and at the moment it needs to happen nobody yet has a stake in the answer. By the time the benefit question is asked, six months later, the pre-deployment state exists only in memory, and memory reliably favours whichever conclusion the person already holds.
A pre and post comparison in a live service cannot separate the tool from everything else that changed. Staffing, seasonality, case mix and the attention that comes with being watched all move at the same time.
You will rarely get a control group in an operational deployment, and it is worth knowing what you are giving up. Where possible, stagger the rollout across comparable units so that a later group serves as an imperfect comparison for an earlier one. It is not a randomised trial, and it is considerably better than nothing.
The uncomfortable truth is that many organisations do not baseline because they do not want the answer to be checkable. A benefit case that cannot be falsified is safer for everyone who signed it.
We think that is a false economy, and the cost lands two or three years later when a finance director asks what the AI portfolio has delivered and nobody can produce a defensible figure. That moment is where AI budgets go to die, and it is entirely preventable with a few hours of work at the start.
Our practical advice: make a baseline a condition of funding. No measurement plan, no budget release. It is the single cheapest governance control available, it takes almost no effort to enforce, and it changes the quality of the proposals you receive because teams start thinking about evidence at the design stage rather than the reporting stage.