Evaluating Feature Selection Methods and Feature Contributions for Cardiovascular Disease Risk Prediction
Akhter, S.; Miller, J. H.
Show abstract
BackgroundCardiovascular disease (CVD) remains the foremost contributor to global illness and death, underscoring the critical need for effective tools that can predict risk at early stages to support preventive care and timely clinical decisions. With the growing complexity of healthcare data, machine learning has shown considerable promise in extracting insights that enhance medical decision-making. Nonetheless, the effectiveness and clarity of machine learning models largely rely on the relevance and quality of input features. MethodsIn this work, we explored and compared three distinct feature selection strategies--Alternating Decision Tree (ADT)-based analysis, Cross-Validated Feature Evaluation (CVFE), and Hypergraph-Based Feature Evaluation (HFE)--to isolate the most predictive clinical variables for assessing CVD risk. Our analysis utilized data from the National Health and Nutrition Examination Survey (NHANES), administered by the National Center for Health Statistics under the Centers for Disease Control and Prevention (CDC), encompassing demographic, clinical, laboratory, and survey data collected across the U.S. from August 2021 through August 2023. Distinct sets of features obtained through the selection techniques were used to develop eXtreme Gradient Boosting (XGBoost) models, which were then assessed for predictive effectiveness. To improve clarity and understand the models decision-making, SHapley Additive exPlanations (SHAP) was utilized to interpret the influence of each feature in the top-performing model. ResultsAmong the approaches, the HFE method achieved the most accurate results, reaching 75% accuracy and an AUC of 0.7857, outperforming the alternatives. The most influential predictors identified by the best model included age, total cholesterol, glycohemoglobin levels, systolic blood pressure, smoking history, and a diagnosis of diabetes. The web application, accessible at https://shiny.tricities.wsu.edu/cvdr-prediction/, presents predictive results, probability scores, and a SHAP plot generated from the model trained using the feature set selected by the hypergraph-based approach. ConclusionsThis study highlights the importance of strategic feature selection in refining predictive accuracy and interpretability, offering a practical data-centric approach that could aid clinicians in evaluating cardiovascular risk and tailoring preventive care. Trial registrationNot applicable as this research is not a clinical trial.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 94%
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 94%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 94%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 92%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 94%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 94%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 94%
- Advancing Cardiovascular Disease Diagnosis: A Robust ML Ecosystem Integrating Early Detection, Responsible AI Framework, and Causal Inference 94%
Similar papers in this journal
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 94%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 93%
- A scoping review of fair machine learning techniques when using real-world data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.