Measuring the Quality of AI-Generated Clinical Notes: A Systematic Review and Experimental Benchmark of Evaluation Methods
Dahlberg, A.; Käennemi, T.; Winther-Jensen, T.; Tapiola, O.; Luisto, R.; Puranen, T.; Gordon, M.; Sanmark, E.; Vartiainen, V.
Show abstract
BackgroundHigh-quality clinical documentation is essential for safe, effective care, yet producing it is time consuming and error prone. Large language models (LLMs) can assist with note generation, but clinical adoption is determined by the resulting note quality. However current evaluation practices vary, and their clinical relevance is unclear. Drawing on a multidisciplinary perspective, we examined how quality is assessed and how those assessments align with clinical demands. MethodsWe systematically searched Ovid Medline and Scopus on 10 April 2025 for peer-reviewed studies that used LLMs in generating clinical notes and included an evaluation of the quality of the resulting text. The screening followed PRISMA and the protocol was preregistered in PROSPERO. Data on metrics, and outcomes were synthesised narratively. Based on these findings, we designed an experimental setup to test the most common evaluation metrics and an LLM-as-evaluator, included for its scalability across large test sets. The experiment used synthetic cases with targeted perturbations. FindingsThirty-seven studies were included. The reporting was dominated by lexical overlap metrics, chiefly ROUGE and BLEU. Semantic similarity metrics, such as BERTScore and BLEURT, were less common. A human evaluation was frequent but heterogeneous, with criteria and methods defined using varying degrees of detail; the most common foci were correctness, fluency, and aspects of clinical acceptability. In our experimental setup, lexical overlap metrics detected deletions and modifications but penalised meaning-preserving paraphrases. Semantic metrics and LLM-as-evaluators were more tolerant of paraphrased perturbations, yet remained sensitive to relevant changes, with performance varying by model and language. InterpretationCurrent practice relies on lexical overlap metrics that are useful for cursory checks but insufficient as proxies for quality. We recommend a layered strategy that pairs semantic metrics with LLM-as-evaluator for scalability and includes targeted human adjudication. Broader, safety-focused validation across institutions and languages is needed before routine deployment. FundingBusiness Finland through the GenAID research project. Personal grants listed under Acknowledgments.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
Similar papers in this journal
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 95%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 94%
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 94%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 94%
- AI-Generated Clinical Summaries: Errors and Susceptibility to Speech and Speaker Variability 90%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 90%
Similar papers in this journal
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 95%
- Large Language Models for Zero-Shot Procedure Extraction in Orthopedic Surgery: A Comparative Evaluation 93%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.