Back

An LLM-Based Comparison of Ambient AI Scribes for Clinical Documentation

Jain, J.; Kaan, J.; Jain, S.; Young, A.; Martinez, C.; Kartsonis, W.; Ortiz, C.; Cheng, R.; Jaklitsch, E.; Cherukuri, S.; Qilleri, A.; Tassiopoulos, A.

2025-06-26 primary care research
10.1101/2025.06.24.25330085 medRxiv
Show abstract

Ambient AI scribes have become an increasingly promising option for automating clinical documentation, with dozens of enterprise solutions available. It remains uncertain whether models with domain-specific tuning outperform naive models "out of the box." This study evaluated five commercial AI scribes, alongside a custom solution using the base model of GPT-o1 without fine-tuning, as well as an experienced human scribe, in a series of simulated clinical encounters. Generated notes from these parties were scored by large language models (LLMs) using a rubric assessing completeness, organization, accuracy, complexity handling, conciseness, and adaptability. Our naive solution achieved scores comparable with industry-leading solutions across all rubric dimensions. These findings suggest that the added value of domain-specific training in ambient AI medical scribes may be limited when compared to base foundation models.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.