Large Language Model Symptom Identification from Clinical Text: A Multi-Center Study
McMurry, A. J.; Phelan, D.; Dixon, B. E.; Geva, A.; Gottlieb, D.; Jones, J. R.; Terry, M.; Taylor, D.; Callaway, H. G.; Manoharan, S.; Miller, T.; Mandl, K. D.
Show abstract
Recognition of patient symptoms is core to medicine, research, and public health. We tested four large language models (LLMs) identifying 11 symptoms of infectious respiratory diseases from emergency department notes (N=204). Each LLM outperformed ICD-10-based identification. GPT-4 had highest tested accuracy, F1 score 91.4% vs. 45.1% for ICD-10. GPT-4 performance in an independent validation cohort (N=308) was even higher with an F1 score of 94.0%.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 93%
- Measuring Quality-of-Care in Treatment of Children with Attention-Deficit/Hyperactivity Disorder: A Novel Application of Natural Language Processing 93%
Similar papers in this journal
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 93%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 92%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 92%
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 91%
- Consistency of performance of adverse outcome prediction models for hospitalized COVID-19 patients 89%
Similar papers in this journal
- Development and validation of a clinical risk score to predict SARS-CoV-2 infection in emergency department patients: The CCEDRRN COVID-19 Infection Score (CCIS) 91%
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 90%
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 90%
Similar papers in this journal
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 94%
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 92%
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.