Back

Assessing ChatGPT's Performance in Delineating Uveitis: An analysis of responses to real-world case presentations

Halim, M. S.; Khowaja, A. H.; Fazal, Z. Z.; Jain, T.; Janjua, K.; Khan, A. A.; Tram Tran, A. N.; Sepah, Y. J.

2025-07-06 ophthalmology
10.1101/2025.07.05.25330926 medRxiv
Show abstract

BackgroundIn the world of Artificial Intelligence (AI), Generative Pretrained Transformer-3 (GPT-3), has gained significant popularity for its demonstrated potential in medical education and diagnostics. RationaleWhile AI has shown promising results in healthcare thus far, its understanding of ocular urgencies, particularly uveitis, demands a focused investigation. MethodsThis study explored the application of ChatGPT, a language model derived from GPT-3, in delineating uveitis based on patient presentations and investigations. We analyzed ChatGPTs communication quality through 14 qualitative metrics by computing patient data at four different levels to act as prompts. These included patient history, drug history, examination findings, and clinical investigations. ResultsOur results showed that at the initial prompt, ChatGPTs responses were comprehensive for most (8 out of 14) variables and correct but inadequate for some (3 out of 14) variables in the majority (>50.0%) of uveitis cases. Ethical considerations was the only variable in terms of which responses consistently showed mixed accuracy and outdated data across all prompts in most (95.8%) uveitis cases. Also, none of the ChatGPT responses were completely inaccurate in terms of any variable at any prompt for any uveitis case. ConclusionThe results reveal ChatGPTs strengths and limitations in answering queries for patients with uveitis or its differential diagnosis while emphasizing the indispensable role of physicians in ethical decision-making.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.