A Randomized-Clinical Trial of Two Ambient Artificial Intelligence Scribes: Measuring Documentation Efficiency and Physician Burnout
Lukac, P. J.; Turner, W.; Vangala, S.; Chin, A. T.; Khalili, J.; Shih, Y.-C. T.; Sarkisian, C.; Cheng, E. M.; Mafi, J. N.
Show abstract
ImportanceAmbient artificial intelligence (AI) scribes record patient encounters and generate visit notes almost instantaneously, representing a promising solution to documentation burden and associated physician burnout. Despite swift and widespread adoption of AI scribes, their impacts have not been examined in randomized-clinical trials. ObjectiveTo test the effectiveness of two AI scribes in reducing time spent writing notes and associated burnout in a randomized-clinical trial. DesignParallel three-arm pragmatic randomized-clinical trial where physicians were assigned 1:1:1 via covariate-constrained randomization (balancing on time-in-note, baseline burnout score, and clinic days /week) to either one of two AI scribe applications--Microsoft DAX or Nabla--or a usual-care control group from 11/4/2024-1/3/2025. SettingA large academic health system in California. Participants313 outpatient physicians were recruited based on leadership referrals and department-wide emails. 238 participants representing 14 specialties qualified. InterventionIntervention-arm physicians gained access to an AI scribe for two months. Main Outcomes and MeasuresThe primary outcome was change from baseline log writing time-in-note. Secondary outcomes measured by surveys included Mini-Z 2.0, 4-item physician task load (TL), and Professional Fulfillment Index-Work Exhaustion (PFI-WE) scores to evaluate aspects of burnout, work environment, and stress, as well as targeted questions addressing safety and accuracy. ResultsDAX was used in 33.5% of 24,696 visits; Nabla was used in 29.5% of 23,653 visits. Nabla users experienced a 9.5% [95% CI:-17.2%,-1.8%] (p=.02) decrease in time-in-note versus the control group and a 7.8% [-15.5%,-0.1%] (p=.05) decrease versus DAX users, while DAX users exhibited no significant change versus control (-1.7% [-9.4%,+5.9%]; p=.66). Total Mini-Z, scaled 10-50 with higher scores indicating improvement, increased with users of any scribe (+2.76 [+1.41,+4.10]; p<.001). Reductions in TL (scale 0-400, TL=-35.8 [-63.7,-7.9]; p=.01) and work exhaustion (scale 0-4, PFI-WE=-0.27 [-0.48,-0.07]; p=.01) were seen with users of any scribe. One Grade 1 (mild) adverse event was reported, while clinically-significant inaccuracies were noted "occasionally" on 5-point Likert questions (DAX 2.7 [2.4-3.0] vs. Nabla 2.8 [2.6-3.0]; p=.68). Conclusion and RelevanceUse of Nabla reduced time-in-note, while use of any scribe led to modest improvements in physician burnout, work exhaustion, and task load. Performance was remarkably similar across two distinct vendor platforms, and occasional inaccuracies observed in either scribe require ongoing physician vigilance. Trial RegistrationClinicalTrials.gov Identifier: NCT06792890
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
- Feasibility characteristics of wrist-worn fitness trackers in health status monitoring for post-COVID patients in remote and rural areas 92%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 94%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
Similar papers in this journal
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 94%
- The Validity of the Parsley Symptom Index: an e-PROM designed for Telehealth 93%
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.