Back

Large language models outperform traditional structured data-based approaches in identifying immunosuppressed patients

Guggilla, V.; Kang, M.; Bak, M. J.; Tran, S. D.; Pawlowski, A.; Nannapaneni, P.; Rasmussen, L. V.; Schneider, D.; Donnelly, H. K.; Agrawal, A.; Liebovitz, D.; Misharin, A. V.; Budinger, G. S.; Wunderink, R. G.; Walunas, T. L.; Gao, C. A.; The NU SCRIPT Study Investigators,

2025-01-17 health informatics
10.1101/2025.01.16.25320564 medRxiv
Show abstract

Identifying immunosuppressed patients using structured data can be challenging. Large language models effectively extract structured concepts from unstructured clinical text. Here we show that GPT-4o outperforms traditional approaches in identifying immunosuppressive conditions and medication use by processing hospital admission notes. We also demonstrate the extensibility of our approach in an external dataset. Cost-effective models like GPT-4o mini and Llama 3.1 also perform well, but not as well as GPT-4o.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.