Identifying anaphylaxis using weakly-supervised prediction models and natural language processing
Williamson, B. D.; Cronkite, D. J.; Yu, O.; Ramaprasan, A.; Fuller, S.; Covey, J.; Kiniry, E.; Park, D.; Winter, R.; Whitaker, J.; McLemore, M. F.; Wittayanukorn, S.; Stojanovic, D.; Zhao, Y.; Dutcher, S.; Carrell, D. S.; Jackson, L. A.; Nelson, J. C.; Smith, J. C.
Show abstract
Objectives Scalable computable phenotyping algorithms are critical for conducting high-throughput disease-outcome research in large, distributed-data electronic health record (EHR) and claims data settings. We developed and evaluated a claims- and EHR-based computable phenotyping algorithm for anaphylaxis, a rare acute condition that is challenging to accurately identify using claims data alone. Materials and Methods Potential anaphylaxis events came from two healthcare systems (Kaiser Permanente Washington [KPWA] and Vanderbilt University Medical Center [VUMC]). We engineered features from clinical text using automated natural language processing (NLP) methods. We then developed a phenotyping algorithm using four NLP- and diagnosis code-based silver labels (proxies for the gold-standard labels). Gold-standard abstracted outcomes were used to evaluate algorithm performance. Results The largest area under the receiver operating characteristic curve (AUC) was 0.931 for an NLP-based silver-label model at KPWA. Depending on the model and healthcare system site, positive predictive value (PPV) and sensitivity at the threshold of predicted probability that maximized F1 score ranged from 0.52 to 0.77 (PPV) and 0.78 to 1 (sensitivity). Discussion NLP-based silver-label models had large AUC at KPWA but not at VUMC. This may be because clinical text at KPWA is only available for outpatient encounters and secure messaging. High sensitivity for identifying anaphylaxis can be obtained using our best-performing models. Conclusion The best-performing models had better PPV and sensitivity tradeoffs than prior bespoke anaphylaxis models with costly, manually curated features. The simplicity of the approach compared to traditional phenotyping methods allows it to be deployed easily at multiple health care systems.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 91%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 90%
- Safety Monitoring of Bivalent COVID-19 mRNA Vaccines Among Recipients 6 months and Older in the United States 90%
Similar papers in this journal
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 94%
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 92%
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 92%
Similar papers in this journal
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 93%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 93%
Similar papers in this journal
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 92%
- Effectiveness of mRNA-1273 against SARS-CoV-2 omicron and delta variants 90%
- Immune Memory Response After a Booster Injection of mRNA-1273 for Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2) 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.