EndoGPT: A Proof-of-concept Large Language Model Based Assistant for the Management of Thyroid Nodules
Shah, M.; Kuo, E. J.; Kuo, J. H.; Hsu, S.; McManus, C.; Liou, R.; Lee, J. A.; Sathe, T. S.
Show abstract
Large language models (LLMs) are increasingly being explored for their potential to simulate clinical reasoning. Here, we demonstrate our initial experience using the GPT-4o LLM along with prompt engineering and knowledge retrieval to develop EndoGPT, a clinical decision support tool for the management of thyroid nodules. In a pilot study of 50 cases, EndoGPT demonstrated an 83% concordance rate with expert surgeons assessments and plans. The highest concordance was in diagnosis (93%), followed by the need for an operation (82%) and type of operation (69%). This work suggests that LLM-based assistants may play a useful role in assisting clinicians in the future.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 90%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 90%
Similar papers in this journal
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 89%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 89%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 87%
Similar papers in this journal
- Real-world, Prospective, and Multicenter Validation of a microRNA-based Thyroid Molecular Classifier 91%
- Annotation-free multi-organ anomaly detection in abdominal CT using free-text radiology reports: A multi-center retrospective study 90%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 88%
Similar papers in this journal
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 91%
- Expert Surgeons and Deep Learning Models Can Predict the Outcome of Surgical Hemorrhage from One Minute of Video 90%
- Large Language Models for Zero-Shot Procedure Extraction in Orthopedic Surgery: A Comparative Evaluation 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.