Back

Evaluation of ChatGPT's Pathology Knowledge using Board-Style Questions

Geetha, S. D.; Khan, A.; Khan, A.; Kannadath, B. S.; Vitkovski, T.

2023-10-03 pathology
10.1101/2023.10.01.23296400 medRxiv
Show abstract

ObjectivesChatGPT is an artificial intelligence (AI) chatbot developed by OpenAI. Its extensive knowledge and unique interactive capabilities enable it to be utilized in various innovative ways in the medical field such as writing clinical notes, simplifying radiology reports. Through this study we aim to analyze its pathology knowledge to advocate its role in transforming pathology education. MethodsAmerican Society for Clinical Pathology (ASCP) Resident Question bank (RQB) 2022 was used to test ChatGPT v4. Practice tests were created in each sub-category and were answered based on the input provided by ChatGPT. Questions that required interpretation of images were excluded. ChatGPTs performance was analyzed and compared with the average peer performance. ResultsThe overall performance of ChatGPT was 56.98%, lower than that of the average peer performance of 62.81%. ChatGPT performed better on clinical pathology (60.42%) than anatomic pathology (54.94%). Furthermore, its performance was better on easy questions (68.47%) compared to intermediate (52.88%) and difficult questions (37.21%). ConclusionsChatGPT has the potential to be a valuable resource in pathology education if trained on a larger, specialized medical dataset. Relying on it solely for the purpose of pathology training should be with caution, in its current form. Key pointsO_LIChatGPT is an AI chatbot, that has gained tremendous popularity in multiple industries, including healthcare. We aim to understand its role in revolutionizing pathology education. C_LIO_LIWe found that ChatGPTs overall performance in Pathology Practice Tests were lower than that expected from an AI tool, furthermore its performance was subpar compared to pathology residents in training. C_LIO_LIIn its current form ChatGPT is not a reliable tool for pathology education, but with further refinement and training it has the potential of being a learning asset. C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.