Back

Large Language Model Symptom Identification from Clinical Text: A Multi-Center Study

McMurry, A. J.; Phelan, D.; Dixon, B. E.; Geva, A.; Gottlieb, D.; Jones, J. R.; Terry, M.; Taylor, D.; Callaway, H. G.; Manoharan, S.; Miller, T.; Mandl, K. D.

2024-12-17 health economics
10.1101/2024.12.16.24319044 medRxiv
Show abstract

Recognition of patient symptoms is core to medicine, research, and public health. We tested four large language models (LLMs) identifying 11 symptoms of infectious respiratory diseases from emergency department notes (N=204). Each LLM outperformed ICD-10-based identification. GPT-4 had highest tested accuracy, F1 score 91.4% vs. 45.1% for ICD-10. GPT-4 performance in an independent validation cohort (N=308) was even higher with an F1 score of 94.0%.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.