← Back to Articles
The short answer

A prospective quality improvement study at Stanford Health Care, published in the Journal of the American Medical Informatics Association in December 2024, found statistically significant improvements after ambient AI scribe deployment: burnout fell 1.94 points (p<.001, Cohen's d 0.86), task load fell 24.42 points (p<.001, d 1.28), and usability rose 10.9 points (p<.001). Median perceived time saving was 20 minutes per half day of clinic. The study enrolled 48 physicians with 38 in the paired analysis, ran October 2023 to January 2024, and had no control group.

What the study found

Stanford University School of Medicine ran a prospective quality improvement study of ambient AI scribes between October 2023 and January 2024, published in the Journal of the American Medical Informatics Association in December 2024. Forty-eight physicians enrolled, with 38 in the paired pre and post analysis.

MeasureBeforeAfterChange
Burnout (work exhaustion)7.475.53−1.94, p<.001, d=0.86
Task load68.9844.56−24.42, p<.001, d=1.28
Usability (SUS)58.0969.01+10.9, p<.001, d=0.69

Median perceived time saving was 20 minutes per half day of clinic. Those are large effect sizes by the standards of workplace interventions.

What it does not establish

Three limits, none of which are hidden by the authors.

  • No control group. A prospective QI study with pre and post measurement cannot separate the tool's effect from everything else that changed over four months, including the attention that comes with being in a pilot.
  • Small and self-selected. Forty-eight physicians who volunteered for an AI scribe pilot are not a random sample of clinicians.
  • Perceived, not measured, time. The 20 minutes is what physicians reported experiencing. Studies that measure documentation time directly have generally found smaller reductions, in the range of 20% to 30%.

Later work has also found that benefits concentrate among clinicians who were coached on how to work with the tool, and that use is inconsistent without that support.

Our view

We are broadly positive on ambient documentation, and we think the way it is being sold sets institutions up for disappointment.

The honest reading of the evidence is that these tools reduce the felt burden of documentation more reliably than they reduce the clock time it consumes. That is not a criticism. Burnout is the outcome most health systems say they care about, and an intervention that moves it is valuable even if the minutes saved are modest. But a business case built on recovered clinical hours, priced against a headcount reduction, is resting on the weaker of the two findings.

The finding we would act on hardest is the coaching one. The organisations getting the most from these tools are the ones that treated deployment as a change in how clinicians work, with training attached, rather than as a licence to switch on. That is an unglamorous conclusion and it is where the variance lives.

Our practical advice: measure your own baseline before deployment, on both dimensions. Documentation time from the audit log, and burnout from a validated instrument. Vendors will supply the second for free. Insist on the first.

Sources

  1. Journal of the American Medical Informatics Association, "Ambient artificial intelligence scribes: physician burnout and perspectives on usability and documentation burden", Stanford University School of Medicine, published 5 December 2024. pmc.ncbi.nlm.nih.gov

Reading Evidence Before Buying

We train teams to evaluate the studies vendors put in front of them.

Start a Conversation →