On the limitations of large language models in clinical diagnosis
Reese, J.; Danis, D.; Caufield, J. H.; Casiraghi, E.; Valentini, G.; Mungall, C. J.; Robinson, P. N.
Show abstract
ObjectiveLarge Language Models such as GPT-4 previously have been applied to differential diagnostic challenges based on published case reports. Published case reports have a sophisticated narrative style that is not readily available from typical electronic health records (EHR). Furthermore, even if such a narrative were available in EHRs, privacy requirements would preclude sending it outside the hospital firewall. We therefore tested a method for parsing clinical texts to extract ontology terms and programmatically generating prompts that by design are free of protected health information. Materials and MethodsWe investigated different methods to prepare prompts from 75 recently published case reports. We transformed the original narratives by extracting structured terms representing phenotypic abnormalities, comorbidities, treatments, and laboratory tests and creating prompts programmatically. ResultsPerformance of all of these approaches was modest, with the correct diagnosis ranked first in only 5.3-17.6% of cases. The performance of the prompts created from structured data was substantially worse than that of the original narrative texts, even if additional information was added following manual review of term extraction. Moreover, different versions of GPT-4 demonstrated substantially different performance on this task. DiscussionThe sensitivity of the performance to the form of the prompt and the instability of results over two GPT-4 versions represent important current limitations to the use of GPT-4 to support diagnosis in real-life clinical settings. ConclusionResearch is needed to identify the best methods for creating prompts from typically available clinical data to support differential diagnostics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
Similar papers in this journal
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 93%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 92%
- Automated Identification of Patients with Immune-related Adverse Events from Clinical Notes using Word embedding and Machine Learning 92%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 91%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.