A Comparison of Manual and Automated Approaches to Developing Computable Algorithms for Identifying Acute Pancreatitis
Bann, M. A.; Carrell, D. S.; Gruber, S.; Heagerty, P. J.; Williamson, B. D.; Nelson, J. C.; Hazlehurst, B.; Felcher, A.; Nyongesa, D. B.; Slaughter, M. T.; Sapp, D. S.; Cronkite, D. J.; Ball, R.; Floyd, J. S.
Show abstract
Objective: Clinical phenotyping methods that rely on clinical and informatics expertise can be time-intensive and costly. We tested both manual and highly automated approaches using electronic health record (EHR) data to identify an FDA Sentinel Initiative health outcome of interest, acute pancreatitis. Materials and Methods: We trained and evaluated machine learning algorithms using EHR data with two approaches: a custom approach that included manually curated features and trained on outcomes data validated with medical record review, and a highly automated approach that greatly simplifies and automates feature engineering and relies on low-cost silver-standard outcomes for model training. Results: Custom algorithms using manually curated structured claims data discriminated cases from non-cases with a high degree of accuracy (cv-AUC 0.89 [95%CI 0.84-0.94]); the inclusion of natural language processing (NLP)-derived covariates from clinical notes increased performance slightly (cv-AUC 0.91[95%CI 0.86-0.97]). The automated algorithm trained on the outcome count of diagnosis codes performed less well (AUC 0.80 [95% CI 0.75-0.85]) but improved using maximum lipase value as an outcome (AUC 0.88 [95% CI 0.84-0.92]). At a positive predictive value of 90%, the custom algorithm had a sensitivity of 92%, the automated algorithm trained on diagnosis code count had a sensitivity of 45%, and the automated algorithm trained on maximum lipase value had a sensitivity of 84%. However, a prediction rule derived by clinicians during chart review was nearly as accurate (maximum lipase value [≥] 3 times upper limit of normal; AUC 0.86, PPV 85%, sensitivity 92%). Discussion: Machine learning algorithms with manually curated structured data and NLP features trained on validated outcomes data successfully identified validated events. Use of an outcome in the automated model based on specific phenotype knowledge (maximum lipase value) allowed for performance similar to the custom model and with considerably less resources.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Approaches for Electronic Health Records Phenotyping: A Methodical Review 93%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 92%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
Similar papers in this journal
- Clinical Study Applying Machine Learning to Detect a Rare Disease: Results and Lessons Learned 92%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 91%
- Modeling physician variability to prioritize relevant medical record information 91%
Similar papers in this journal
- Developing deep learning-based strategies to predict the risk of hepatocellular carcinoma among patients with nonalcoholic fatty liver disease from electronic health records 92%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 91%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 91%
Similar papers in this journal
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 92%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 92%
- CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization 91%
Similar papers in this journal
- Augmented Intelligence with Natural Language Processing Applied to Electronic Health Records is Useful for Identifying Patients with Non-Alcoholic Fatty Liver Disease at Risk for Disease Progression 92%
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 91%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.