Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study
Cotte, F.; Schmude, M.; Bode, P.; Suliman, O.; Dias Lourenco, F.; Paiva Pereira, M.; Kini, N.; Hartenstein, V.; Muscoloni, A.; Stroux, L.; Hertz, V.; Koehler, S.; Morelli, V.; Hoffmann, H.; Engerer, P.; Gilbert, S.; Gray, K.; Mehrali, T.; Seemann Monteiro, M.; Flores, P. D.
Show abstract
BackgroundDiagnostic decision support systems (DDSS) are increasingly deployed at scale, yet their diagnostic accuracy is insufficiently monitored once integrated into care. Traditional post-market surveillance relies on clinician review, which is costly, slow, and difficult to sustain. Large language models (LLMs) may offer a scalable and potentially automatable solution, but their performance in real-world monitoring remains unknown. MethodsWe conducted a diagnostic accuracy substudy within ESSENCE, a prospective evaluation of Ada Healths DDSS integrated into Portugals largest private healthcare network. Clinical notes and ICD-10 diagnoses from 498 encounters were anonymised and classified using a filter-map-match framework. Manual clinician review served as the reference standard. We compared eligibility classification and condition mapping between clinicians and GPT-4.1 and GPT-5, and assessed diagnostic accuracy of two DDSS versions using both reference sets. FindingsManual review classified 385 of 498 encounters (77{middle dot}3%) as eligible for diagnostic comparison. GPT-5 reproduced these classifications with 84{middle dot}7% accuracy ({kappa}=0{middle dot}57), showing high sensitivity but only moderate specificity. Among 347 encounters judged eligible by both approaches, GPT-5 exactly matched clinician-assigned diagnoses in 93{middle dot}6% and proposed clinically plausible alternatives in 3{middle dot}5%. Diagnostic accuracy estimates based on manual versus GPT-5 mappings were statistically indistinguishable at Top-1 and Top-3 across the full analyzable sets, with one significant difference at Top-5. In the overlapping 346 cases, no statistical differences were observed. Across both reference sets, the experimental DDSS version outperformed the original only at the Top-5 threshold. InterpretationLLMs can reproduce clinician review of real-world diagnostic encounters with close agreement. While GPT-5 performed comparably to clinicians for condition mapping, the eligibility filtering step - deciding which encounters should enter the diagnostic-accuracy analysis - remains the main source of divergence and is the priority for improvement. Embedding such approaches into health systems could enable automated and continuous performance and safety monitoring and support regulatory compliance. Broader evaluations across diverse care settings are needed to establish generalisability and equity impact. FundingGerman Federal Ministry of Education and Research (NextGenerationEU, PATH project).
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 94%
- A data management system for precision medicine 92%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 92%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 94%
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 92%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.