← Back to Articles
The short answer

Before deployment, capture four things: the time the current process actually takes measured from system logs rather than estimated, the current error or rework rate, a validated measure of staff experience for the people whose work changes, and the volume the process handles. Without a pre-deployment baseline, any later benefit claim is an estimate arguing with a budget, and the organisation cannot tell a working deployment from a popular one.

Why this gets skipped

Baselining is boring, it delays the exciting part, and at the moment it needs to happen nobody yet has a stake in the answer. By the time the benefit question is asked, six months later, the pre-deployment state exists only in memory, and memory reliably favours whichever conclusion the person already holds.

The four measures

  • Process time, measured. From audit logs or timestamps, not from asking people how long it takes. Self-reported process time is consistently wrong, usually in the direction of the person's frustration with the task.
  • Error and rework rate. How often the current process produces something that has to be corrected. If the AI is faster and wrong more often, a time-only measure will show a win.
  • Staff experience, on a validated instrument. For the people whose work actually changes. Burnout and task load instruments exist and are free to use, and they give you a defensible number rather than an anecdote.
  • Volume. How many times the process runs. A 30% improvement on something that happens twice a week is not a business case.

The comparison problem

A pre and post comparison in a live service cannot separate the tool from everything else that changed. Staffing, seasonality, case mix and the attention that comes with being watched all move at the same time.

You will rarely get a control group in an operational deployment, and it is worth knowing what you are giving up. Where possible, stagger the rollout across comparable units so that a later group serves as an imperfect comparison for an earlier one. It is not a randomised trial, and it is considerably better than nothing.

Our view

The uncomfortable truth is that many organisations do not baseline because they do not want the answer to be checkable. A benefit case that cannot be falsified is safer for everyone who signed it.

We think that is a false economy, and the cost lands two or three years later when a finance director asks what the AI portfolio has delivered and nobody can produce a defensible figure. That moment is where AI budgets go to die, and it is entirely preventable with a few hours of work at the start.

Our practical advice: make a baseline a condition of funding. No measurement plan, no budget release. It is the single cheapest governance control available, it takes almost no effort to enforce, and it changes the quality of the proposals you receive because teams start thinking about evidence at the design stage rather than the reporting stage.

Sources

  1. Ministry of Health Singapore, AIHGle 2.0 deployers' toolkit, on internal governance for new adopters. go.gov.sg/aihgle

Knowing What Good Looks Like

We help teams define the measures before the technology arrives.

Start a Conversation →