Evaluation of the Diagnostic Accuracy of GPT-4 in Five Thousand Rare Disease Cases
Reese, J. T.; Chimirri, L.; Danis, D.; Caufield, J. H.; Wissink, K. W.; Casiraghi, E.; Valentini, G.; Haendel, M. A.; Mungall, C. J.; Robinson, P. N.
Show abstract
Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5,213 case reports using the Phenopacket Schema, the Human Phenotype Ontology and Mondo disease ontology. Prompts generated from each phenopacket were sent to seven LLMs, including four generalist models and three LLMs specialized for medical applications. The same phenopackets were used as input to a widely used diagnostic tool, Exomiser, in phenotype-only mode. The best LLM ranked the correct diagnosis first in 23.6% of cases, whereas Exomiser did so in 35.5% of cases. While the performance of LLMs for supporting differential diagnosis has been improving, it has not reached the level of commonly used traditional bioinformatics tools. Future research is needed to determine the best approach to incorporate LLMs into diagnostic pipelines.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 93%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 92%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Developing and Evaluating Pediatric Phecodes (Peds-Phecodes) for High-Throughput Phenotyping Using Electronic Health Records 92%
Similar papers in this journal
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 94%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 93%
- Development and validation of AI-based pre-screening of large bowel biopsies 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.