Evaluating the Performance of Artificial Intelligence in Generating Differential Diagnoses for Infectious Diseases Cases: A Comparative Study of Large Language Models
Mondal, A.; Karad, R. K.; Bhattacharjee, B.; Saha, B.
Show abstract
BackgroundArtificial Intelligence (AI) has potential to transform healthcare including the field of infectious diseases diagnostics. This study assesses the capability of three large language models (LLMs), GPT 4, Llama 3, and Gemini 1.5 to generate differential diagnoses, comparing their outputs against those of medical experts to evaluate AIs potential in augmenting clinical decision-making. MethodsThis study evaluates the differential diagnosis capabilities of three LLMs, GPT 4, Llama 3, and Gemini 1.5, using 50 simulated infectious disease cases. The cases were diverse, complex, and reflective of common clinical scenarios, including detailed histories, symptoms, lab results, and imaging findings. Each model received standardized case information and produced differential diagnoses, which were then compared to reference differential diagnosis lists created by medical experts. The analysis utilized the Jaccard index and Kendalls Tau to assess similarity and order accuracy, summarizing findings with mean, standard deviation, and combined p-values. ResultsThe mean numbers of differential diagnoses generated by GPT 4, Llama 3, and Gemini 1.5 were 6.22, 5.06, and 10.02 respectively which was significantly different (p<0.001) from the medical experts. The mean Jac-card index of GPT 4, Llama 3, and Gemini 1.5 were 0.3, 0.21, and 0.24 while the mean Kendalls Tau were 0.4, 0.7, and 0.33 respectively. The combined p-value of GPT 4, Llama 3, and Gemini 1.5 were 1, 1, 0.979 respectively indicating no significant association between the differential diagnosis generated by the LLMs and the medical experts. ConclusionAlthough LLMs like GPT 4, Llama 3, and Gemini 1.5 exhibit varying effectiveness, none align significantly with expert-level diagnostic accuracy, emphasizing the need for further development and refinement. The findings highlight the importance of rigorous validation, ethical considerations, and seamless integration into clinical workflows to ensure AI tools enhance healthcare delivery and patient outcomes effectively.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 93%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 95%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 94%
- Comparison of OneChoice(R) AI-based clinical decision support recommendations with infectious disease specialists and non-specialists for bacteremia treatment in Lima, Peru 93%
Similar papers in this journal
- Machine Learning Interpretability Methods to Characterize the Importance of Hematologic Biomarkers in Prognosticating Patients with Suspected Infection 93%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 93%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 92%
Similar papers in this journal
- Explanation of Hand, Foot, and Mouth Disease Cases in Japan Using Google Trends Before and During the COVID-19: Infodemiology Study 90%
- Guidelines on reporting and assessing mathematical models for infectious disease dynamics: A scoping review 90%
- Estimating effects of intervention measures on COVID-19 outbreak in Wuhan taking account of improving diagnostic capabilities using a modelling approach 89%
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 94%
- Large language models for generating medical examinations: systematic review 94%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.