Hazard-aware adaptations bridge the generalization gap in large language models: a nationwide study
Wu, J.; Conover, S.; Su, C.; Corrigan, J.; Culnan, J.; Liu, Y.; Kelley, M.; Do, N.; Arya, S.; Sox-Harris, A.; Langlotz, C.; Wiener, R.; Branch-Elliman, W.; Han, S.; Fillmore, N.
Show abstract
Despite growing excitement in deploying large language models (LLMs) for healthcare, most machine learning studies show success on the same few limited public data sources. It is unclear if and how most results generalize to real-world clinical settings. To measure this gap and shorten it, we analyzed protected notes from over 100 Veterans Affairs (VA) sites, focusing on extracting smoking history--a persistent and clinically impactful problem in natural language processing (NLP). Here we applied adaptation techniques to an LLM over two institutional datasets, a popular public dataset (MIMIC-III) and our VA one, across five smoking history NLP tasks of varying complexity. We demonstrate that adapted prompts, engineered to address observed errors, achieve better generalizability across institutions compared to zero-shot prompts. We analyzed 2,955 notes and LLM outputs to codify errors in a hazard framework, identifying whether error frequency differences between institutions stemmed from generalization failures or inherent data differences. While overall accuracy with the adapted prompt was similar between institutions (macro-F1=0.86 in VA, 0.85 in MIMIC), hazard distributions varied significantly. In some cases, a dataset had more errors in a specific category due to a higher prevalence of the associated hazard, such as templated information in VA notes (padj=0.004). However, when task-specific requirements conflicted with pre-trained model behavior, errors in the untrained institution were more frequent despite similar hazard prevalence (padj=0.007), showing a limit of LLM generalizability. As a potential clinical application, our adapted LLM system identified lung cancer screening eligibility in 59% of Veterans who later developed the disease, compared to 8% with current national VA tools. Our results demonstrate LLM generalizability on real-world, national patient data while identifying hazards to address for improved performance and broader applicability.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 94%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.