An LLM-Based Comparison of Ambient AI Scribes for Clinical Documentation
Jain, J.; Kaan, J.; Jain, S.; Young, A.; Martinez, C.; Kartsonis, W.; Ortiz, C.; Cheng, R.; Jaklitsch, E.; Cherukuri, S.; Qilleri, A.; Tassiopoulos, A.
Show abstract
Ambient AI scribes have become an increasingly promising option for automating clinical documentation, with dozens of enterprise solutions available. It remains uncertain whether models with domain-specific tuning outperform naive models "out of the box." This study evaluated five commercial AI scribes, alongside a custom solution using the base model of GPT-o1 without fine-tuning, as well as an experienced human scribe, in a series of simulated clinical encounters. Generated notes from these parties were scored by large language models (LLMs) using a rubric assessing completeness, organization, accuracy, complexity handling, conciseness, and adaptability. Our naive solution achieved scores comparable with industry-leading solutions across all rubric dimensions. These findings suggest that the added value of domain-specific training in ambient AI medical scribes may be limited when compared to base foundation models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- AI-Generated Clinical Summaries: Errors and Susceptibility to Speech and Speaker Variability 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
Similar papers in this journal
Similar papers in this journal
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.