Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study
Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.
Show abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 93%
- Beyond Metrics to Methods: A Scoping Review of Large Language Models for Detection of Social Drivers of Health in Clinical Notes 92%
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 91%
Similar papers in this journal
- Ambient scribe in general practice: a multi-perspective before-after longitudinal mixed-methods study 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
Similar papers in this journal
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 92%
- Clinical practice guidelines and recommendations in the context of the COVID-19 pandemic: systematic review and critical appraisal 91%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 90%
Similar papers in this journal
- Development of the Tool for Advancing Practice Performance, a practice-level survey to assess primary care structures and processes 92%
- Feasibility trial of a new digital training package to enhance primary care practitioners' communication of clinical empathy and realistic optimism 91%
- Clinical academic research in the time of Corona: a simulation study in England and a call for action 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.