Performance, Generalizability, and Fairness of a Peripheral Artery Disease Detection Model Across Patient Phenotypes and Health Systems
Kallis, K.; Quitevis, C. R.; Ramsis, M.; Kabutey, N.-K.; Conte, M. S.; Rowe, V. L.; Humphries, M. D.; Hernandez-Boussard, T.; B. Malas, M.; Ross, E. G.
Show abstract
Background Peripheral artery disease (PAD) is a major cause of cardiovascular events but remains underdiagnosed. Electronic health record (EHR)-based machine learning models show promise for earlier detection, but developing generalizable and fair models across diverse populations remains challenging. Methods Using the University of California Health Data Warehouse, containing EHR data from five health systems, we identified patients with and without PAD. We used unsupervised clustering to define PAD phenotypes and trained a LightGBM classifier using 14,023 features spanning demographics, comorbidities, medications, laboratory values, healthcare utilization, and diagnosis, procedure, and medication codes. We evaluated performance overall and across demographic groups and phenotypes, and assessed fairness using selection rates and subgroup differences in true- and false-positive rates. Results The study included 33,739 cases and 33,739 matched controls. Clustering identified four phenotypes: patients with limited healthcare documentation (cluster 1), younger patients with severe metabolic disease (cluster 2), patients with a traditional atherosclerotic risk profile (cluster 3), and frail elderly patients with multimorbidity (cluster 4). Overall, the model demonstrated consistent performance across institutions (AUROC 0.76?0.79; AUC-PR 0.76?0.79) with well-calibrated probabilities. Performance was similar across genders, with modest variation by race and age, and was stronger in clusters 2?4. Cluster 2 demonstrated the highest sensitivity (TPR 0.87, 95% CI 0.87?0.88), while cluster 1 showed the lowest performance (TPR 0.40, 95% CI 0.39?0.41). Conclusions The EHR-based PAD detection model demonstrated consistent performance across five health systems. Phenotypic clustering revealed clinically meaningful differences in model performance adding an additional consideration in ML fairness and performance evaluations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Machine Learning Approaches for Electronic Health Records Phenotyping: A Methodical Review 92%
Similar papers in this journal
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
- Response to Polygenic Risk: Results of the MyGeneRank Mobile Application-Based Coronary Artery Disease Study 92%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 92%
Similar papers in this journal
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 93%
- Modeling physician variability to prioritize relevant medical record information 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
Similar papers in this journal
- Predicting 30-Day and 1-Year Mortality in Heart Failure with Preserved Ejection Fraction (HFpEF) 93%
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 93%
- Enhanced machine learning and hybrid ensemble approaches for coronary heart disease prediction 93%
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 92%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 91%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.