Evaluating the accuracy and reliability of large language models in assisting with pediatric differential diagnoses: A multicenter diagnostic study
Mansoor, M. A.; Ibrahim, A. F.; Grindem, D. J.; Baig, A.
10.1101/2024.08.09.24311777 medRxivShow abstract
ImportanceLarge language models, such as GPT-3, have shown potential in assisting with clinical decision-making, but their accuracy and reliability in pediatric differential diagnosis in rural healthcare settings remain underexplored. ObjectiveEvaluate the performance of a fine-tuned GPT-3 model in assisting with pediatric differential diagnosis in rural healthcare settings and compare its accuracy to human physicians. MethodsRetrospective cohort study using data from a multicenter rural pediatric healthcare organization in Central Louisiana serving approximately 15,000 patients. Data from 500 pediatric patient encounters (age range: 0-18 years) between March 2023 and January 2024 were collected and split into training (70%, n=350) and testing (30%, n=150) sets. InterventionsGPT-3 model (DaVinci version) fine-tuned using OpenAI API on training data for ten epochs. Main Outcomes and MeasuresAccuracy of fine-tuned GPT-3 model in generating differential diagnoses, evaluated using sensitivity, specificity, precision, F1 score, and overall accuracy. The models performance was compared to human physicians on the testing set. ResultsThe fine-tuned GPT-3 model achieved an accuracy of 87% (131/150) on the testing set, with a sensitivity of 85%, specificity of 90%, precision of 88%, and F1 score of 0.87. The models performance was comparable to human physicians (accuracy 91%; P = .47). Conclusions and RelevanceThe fine-tuned GPT-3 model demonstrated high accuracy and reliability in assisting with pediatric differential diagnosis, with performance comparable to human physicians. Large language models could be valuable tools for supporting clinical decision-making in resource-constrained environments. Further research should explore implementation in various clinical workflows.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- External validation of a paediatric SMART triage model for use in resource limited facilities 94%
- Machine Learning for Paediatric Related Decision Support in Emergency Care - A UK and Ireland Network Survey Study 94%
- ePOCT+ and the medAL-suite: Development of an electronic clinical decision support algorithm and digital platform for pediatric outpatients in low- and middle-income countries 93%
Similar papers in this journal
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Clinical code sets and the problem of redundancy in code set repositories 92%
Similar papers in this journal
- A Retrospective Cohort Study of COVID-19 among Children in Fulton County, Georgia, March 2020 - June 2021 89%
- How young people experienced Long COVID services: a qualitative analysis 89%
- Towards uniform recognition of child abuse in the Netherlands: implementing the Screening instrument for Child Abuse and Neglect (SCAN) 88%
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 91%
Similar papers in this journal
- Revisiting the use and effectiveness of patient-held records in rural Malawi 91%
- Predictive accuracy of computer-aided versions of the on-admission National Early Warning Score in estimating the risk of COVID-19 for unplanned admission to hospital: a retrospective development and validation study 91%
- Assessing Language Difficulties in Health Facilities in Malawi 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.