Evaluation of ChatGPT's Usefulness and Accuracy in Diagnostic Surgical Pathology.
Guastafierro, V.; Corbitt, D. N.; Bressan, A.; Fernandes, B.; Mintemur, O.; Magnoli, F.; Ronchi, S.; La Rosa, S.; Uccella, S.; Renne, S. L.
Show abstract
ChatGPT is an artificial intelligence capable of processing and generating human-like language. ChatGPTs role within clinical patient care and medical education has been explored; however, assessment of its potential in supporting histopathological diagnosis is lacking. In this study, we assessed ChatGPTs reliability in addressing pathology-related diagnostic questions across 10 subspecialties, as well as its ability to provide scientific references. We created five clinico-pathological scenarios for each subspecialty, posed to ChatGPT as open-ended or multiple-choice questions. Each question either asked for scientific references or not. Outputs were assessed by six pathologists according to: 1) usefulness in supporting the diagnosis and 2) absolute number of errors. All references were manually verified. We used directed acyclic graphs and structural causal models to determine the effect of each scenario type, field, question modality and pathologist evaluation. Overall, we yielded 894 evaluations. ChatGPT provided useful answers in 62.2% of cases. 32.1% of outputs contained no errors, while the remaining contained at least one error (maximum 18). ChatGPT provided 214 bibliographic references: 70.1% were correct, 12.1% were inaccurate and 17.8% did not correspond to a publication. Scenario variability had the greatest impact on ratings, followed by prompting strategy. Finally, latent knowledge across the fields showed minimal variation. In conclusion, ChatGPT provided useful responses in one-third of cases, but the number of errors and variability highlight that it is not yet adequate for everyday diagnostic practice and should be used with discretion as a support tool. The lack of thoroughness in providing references also suggests caution should be employed even when used as a self-learning tool. It is essential to recognize the irreplaceable role of human experts in synthesizing images, clinical data and experience for the intricate task of histopathological diagnosis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 94%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 94%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 93%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 93%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 93%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 94%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 93%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 91%
Similar papers in this journal
- Development of an Interactive Web Dashboard to Facilitate the Reexamination of Pathology Reports for Instances of Underbilling of CPT Codes 97%
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 93%
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.