The most-cited study on AI scribes reports strong effects on burnout and documentation load. It also has 48 participants and no control group.
A prospective quality improvement study at Stanford Health Care, published in the Journal of the American Medical Informatics Association in December 2024, found statistically significant improvements after ambient AI scribe deployment: burnout fell 1.94 points (p<.001, Cohen's d 0.86), task load fell 24.42 points (p<.001, d 1.28), and usability rose 10.9 points (p<.001). Median perceived time saving was 20 minutes per half day of clinic. The study enrolled 48 physicians with 38 in the paired analysis, ran October 2023 to January 2024, and had no control group.
Stanford University School of Medicine ran a prospective quality improvement study of ambient AI scribes between October 2023 and January 2024, published in the Journal of the American Medical Informatics Association in December 2024. Forty-eight physicians enrolled, with 38 in the paired pre and post analysis.
| Measure | Before | After | Change |
|---|---|---|---|
| Burnout (work exhaustion) | 7.47 | 5.53 | −1.94, p<.001, d=0.86 |
| Task load | 68.98 | 44.56 | −24.42, p<.001, d=1.28 |
| Usability (SUS) | 58.09 | 69.01 | +10.9, p<.001, d=0.69 |
Median perceived time saving was 20 minutes per half day of clinic. Those are large effect sizes by the standards of workplace interventions.
Three limits, none of which are hidden by the authors.
Later work has also found that benefits concentrate among clinicians who were coached on how to work with the tool, and that use is inconsistent without that support.
We are broadly positive on ambient documentation, and we think the way it is being sold sets institutions up for disappointment.
The honest reading of the evidence is that these tools reduce the felt burden of documentation more reliably than they reduce the clock time it consumes. That is not a criticism. Burnout is the outcome most health systems say they care about, and an intervention that moves it is valuable even if the minutes saved are modest. But a business case built on recovered clinical hours, priced against a headcount reduction, is resting on the weaker of the two findings.
The finding we would act on hardest is the coaching one. The organisations getting the most from these tools are the ones that treated deployment as a change in how clinicians work, with training attached, rather than as a licence to switch on. That is an unglamorous conclusion and it is where the variance lives.
Our practical advice: measure your own baseline before deployment, on both dimensions. Documentation time from the audit log, and burnout from a validated instrument. Vendors will supply the second for free. Insist on the first.