← Back to Articles
The short answer

Before signing, establish six things: what the model was validated on and whether that population resembles yours, how performance is monitored after go-live and who watches it, what the system does when it is uncertain, who is accountable when output is wrong, what happens to your data, and how you exit. A vendor who cannot answer these precisely is telling you something useful.

Why evaluation is where it goes wrong

In most healthcare AI failures we see, the deciding mistake was made before implementation started. The organisation could not tell a substantive answer from a fluent one, so it bought on demo quality and reference customers. Both are weak signals. A good demo shows the system working on data the vendor chose.

Under Singapore's AIHGle 2.0, the deployer carries responsibilities that cannot be transferred to the developer, which makes this conversation part of your governance obligation rather than just good practice.

Validation and population fit

  • What population was this validated on, and how does it compare to ours? Age, ethnicity, comorbidity mix, care setting. A model validated on a North American academic centre may behave differently in a Singapore polyclinic.
  • Show us the performance breakdown by subgroup. Aggregate accuracy hides differential performance. If they have not measured it by subgroup, they cannot tell you it is equitable.
  • What is the false positive and false negative rate at the threshold we would run? Not the best achievable figure. The one at the operating point they are proposing.
  • Has this been independently evaluated, and by whom? Internal validation by the vendor is a starting point, not an endorsement.

Monitoring and drift

  • How will we know if performance degrades after go-live? Populations shift, documentation practices change, upstream systems get upgraded. Ask what would surface a problem, and how fast.
  • Who is responsible for watching that, us or you? This is frequently assumed by both parties and owned by neither.
  • How often is the model updated, and do we get a say? A silent update can change behaviour clinicians have calibrated to.

Failure modes and human oversight

  • What does the system do when it is uncertain? A system that abstains and escalates is safer than one that always produces an answer.
  • What is the worst realistic failure, and what would it look like from the clinician's side? If they have not thought about this, they have not thought about safety.
  • What does a user need to be told to exercise judgment properly? AIHGle 2.0's transparency guidance exists so users can decide how much weight to give an output.

Data, liability and exit

  • Where does our data go, who can access it, and is it used to train your models? Get this in the contract, not the sales conversation.
  • Who is liable when the output is wrong and a patient is harmed? Ask for the answer in writing. Ambiguity here defaults to you.

And the one people forget: what happens if we want to leave in three years. Can you extract your data in a usable format, and what stops working the day the contract ends?

Our view

The purpose of this list is not to catch vendors out. Good vendors welcome these questions, because answering them well is how they differentiate from competitors selling on demo polish. The ones who get uncomfortable are telling you something you need to know.

The harder problem is that asking the question is not the same as evaluating the answer. Any procurement team can read a checklist aloud. Knowing whether a validation study actually supports the claim, or whether a proposed monitoring approach would realistically catch drift, requires enough fluency to ask the follow-up. That is the capability gap we spend most of our time closing, and it is why we would rather train the people in the room than hand over a longer checklist.

One rule we give clients: never let the vendor be the only party in the room who understands the system. If that is the case, you are not negotiating, you are being briefed.

Sources

  1. Ministry of Health Singapore, Artificial Intelligence in Healthcare Guidelines (AIHGle 2.0), including the deployers' toolkit. go.gov.sg/aihgle
  2. Health Sciences Authority, Regulatory Guidelines for Software as Medical Devices. hsa.gov.sg

Sharper Questions, Better Decisions

We train leadership teams to run this conversation themselves.

Start a Conversation →