Assessing ChatGPT's Performance in Delineating Uveitis: An analysis of responses to real-world case presentations
Halim, M. S.; Khowaja, A. H.; Fazal, Z. Z.; Jain, T.; Janjua, K.; Khan, A. A.; Tram Tran, A. N.; Sepah, Y. J.
Show abstract
BackgroundIn the world of Artificial Intelligence (AI), Generative Pretrained Transformer-3 (GPT-3), has gained significant popularity for its demonstrated potential in medical education and diagnostics. RationaleWhile AI has shown promising results in healthcare thus far, its understanding of ocular urgencies, particularly uveitis, demands a focused investigation. MethodsThis study explored the application of ChatGPT, a language model derived from GPT-3, in delineating uveitis based on patient presentations and investigations. We analyzed ChatGPTs communication quality through 14 qualitative metrics by computing patient data at four different levels to act as prompts. These included patient history, drug history, examination findings, and clinical investigations. ResultsOur results showed that at the initial prompt, ChatGPTs responses were comprehensive for most (8 out of 14) variables and correct but inadequate for some (3 out of 14) variables in the majority (>50.0%) of uveitis cases. Ethical considerations was the only variable in terms of which responses consistently showed mixed accuracy and outdated data across all prompts in most (95.8%) uveitis cases. Also, none of the ChatGPT responses were completely inaccurate in terms of any variable at any prompt for any uveitis case. ConclusionThe results reveal ChatGPTs strengths and limitations in answering queries for patients with uveitis or its differential diagnosis while emphasizing the indispensable role of physicians in ethical decision-making.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 95%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 92%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 92%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 96%
- Clinical code sets and the problem of redundancy in code set repositories 93%
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 93%
Similar papers in this journal
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 92%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 92%
- The experiences of 33 national COVID-19 dashboard teams during the first year of the pandemic in the WHO European Region: a qualitative study 91%
Similar papers in this journal
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 95%
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 94%
- Autonomous Screening for Laser Photocoagulation in Fundus Images Using Deep Learning 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.