Evaluation of the Clinical Utility of DxGPT, a GPT-4 Based Large Language Model, through an Analysis of Diagnostic Accuracy and User Experience
Alvarez-Estape, M.; Cano, I.; Pino, R.; Gonzalez Grado, C.; Aldemira-Liz, A.; Gonzalvez-Ortuno, J.; do Olmo, J.; Logrono, J.; Martinez, M.; Mascias, C.; Isla, J.; Martinez Roldan, J.; Launes, C.; Garcia-Cuyas, F.; Esteller-Cucala, P.
Show abstract
ImportanceThe time to accurately diagnose rare pediatric diseases often spans years. Assessing the diagnostic accuracy of an LLM-based tool on real pediatric cases can help reduce this time, providing quicker diagnoses for patients and their families. ObjectiveTo evaluate the clinical utility of DxGPT as a support tool for differential diagnosis of both common and rare diseases. DesignUnicentric descriptive cross-sectional exploratory study. Anonymized data from 50 pediatric patients medical histories, covering common and rare pathologies, were used to generate clinical case notes. Each clinical case included essential data, with some expanded by complementary tests. SettingThis study was conducted at a reference pediatric hospital, Sant Joan de Deu Barcelona Childrens Hospital. ParticipantsA total of 50 clinical cases were diagnosed by 78 volunteer doctors (medical diagnostic team) with varying experience, each reviewing 3 clinical cases. InterventionsEach clinician listed up to five diagnoses per clinical case note. The same was done on the DxGPT web platform, obtaining the Top-5 diagnostic proposals. To evaluate DxGPTs variability, each note was queried three times. Main Outcome(s) and Measure(s)The study mainly focused on comparing diagnostic accuracy, defined as the percentage of cases with the correct diagnosis, between the medical diagnostic team and DxGPT. Other evaluation criteria included qualitative assessments. The medical diagnostic team also completed a survey on their user experience with DxGPT. ResultsTop-5 diagnostic accuracy was 65% for clinicians and 60% for DxGPT, with no significant differences. Accuracies for common diseases were higher (Clinicians: 79%, DxGPT: 71%) than for rare diseases (Clinicians: 50%, DxGPT: 49%). Accuracy increased similarly in both groups with expanded information, but this increase was only stastitically significant in clinicians (simple 52% vs. expanded 69%; p=0.03). DxGPTs response variability affected less than 5% of clinical case notes. A survey of 48 clinicians rated the DxGPT platform 3.9/5 overall, 4.1/5 for usefulness, and 4.5/5 for usability. Conclusions and RelevanceDxGPT showed diagnostic accuracies similar to medical staff from a pediatric hospital, indicating its potential for supporting differential diagnosis in other settings. Clinicians praised its usability and simplicity. These tools could provide new insights for challenging diagnostic cases. Key Points QuestionIs DxGPT, a large language model-based (LLM-based) tool, effective for differential diagnosis support, specifically in the context of a clinical pediatric setting? FindingsIn this unicentric cross-sectional study, diagnostic accuracy, measured as the proportion of clinical cases where any of the five diagnostic options included the correct diagnosis, showed comparable results between clinicians and DxGPT. Top-5 accuracy was 65% for clinicians and 60% for DxGPT. MeaningThese findings highlight the potential of LLM-based tools like DxGPT to support clinicians in making accurate and timely diagnoses, ultimately improving patient care.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of a Large Language Model to Identify Confidential Content in Adolescent Encounter Notes 92%
- Clinical features and burden of post-acute sequelae of SARS-CoV-2 infection in children and adolescents: an exploratory EHR-based cohort study from the RECOVER program 92%
- Acute upper airway disease in children with the omicron (B.1.1.529) variant of SARS-CoV-2: a report from the National COVID Cohort Collaborative (N3C) 89%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Developing and Evaluating Pediatric Phecodes (Peds-Phecodes) for High-Throughput Phenotyping Using Electronic Health Records 93%
- Measuring Quality-of-Care in Treatment of Children with Attention-Deficit/Hyperactivity Disorder: A Novel Application of Natural Language Processing 93%
Similar papers in this journal
- DDIEM: Drug Database for Inborn Errors of Metabolism 90%
- Good communication is critical to supporting people living and working with a rare disease: current rare disease support perceived as inadequate. 89%
- The COVID-19 pandemic impact on continuity of care provision on rare brain diseases and on Ataxia, Dystonia and PKU. A scoping review protocol 89%
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 94%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 91%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.