Combining Machine Learning with Cox models for identifying risk factors for incident post-menopausal breast cancer in the UK Biobank
Liu, X.; Collister, J. A.; Littlejohns, T. J.; Morelli, D.; Clifton, D. A.; Hunter, D. J.; Clifton, L.
Show abstract
1.Breast cancer is the most common cancer in women. A better understanding of risk factors plays a central role in disease prediction and prevention. We aimed to identify potential novel risk factors for breast cancer among post-menopausal women, with pre-specified interest in the role of polygenic risk scores (PRS) for risk prediction. We designed an analysis pipeline combining both machine learning (ML) and classical statistical models with emphasis on necessary statistical considerations (e.g. collinearity, missing data). Extreme gradient boosting (XGBoost) machine with Shapley (SHAP) feature importance measures were used for risk factor discovery among [~]1.7k features in 104,313 post-menopausal women from the UK Biobank cohort. Cox models were constructed subsequently for in-depth investigation. Both PRS were significant risk factors when fitted simultaneously in both ML and Cox models (p < 0.001). ML analyses identified 11 (excluding the two PRS) novel predictors, among which five were confirmed by the Cox models: plasma urea (HR=0.95, 95% CI 0.92-0.98, p < 0.001) and plasma phosphate (HR=0.67, 95% CI 0.52-0.88, p = 0.003) were inversely associated with risk of developing post-menopausal breast cancer, whereas basal metabolic rate (HR=1.15, 95% CI 1.08-1.22, p < 0.001), red blood cell count (HR=1.20, 95% CI 1.08-1.34, p = 0.001), and creatinine in urine (HR=1.05, 95% CI 1.01-1.09, p = 0.008) were positively associated. Our final Cox model demonstrated a slight improvement in risk discrimination when adding novel features to a simpler Cox model containing PRS and the established risk factors (Harrells C-index = 0.670 vs 0.665).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The significance of molecular heterogeneity in breast cancer batch correction and dataset integration 94%
- Comparative validation of the BOADICEA and Tyrer-Cuzick breast cancer risk models incorporating classical risk factors and polygenic risk in a population-based prospective cohort 94%
- Clustering of HR+/HER2- breast cancer in an Asian cohort is driven by immune phenotypes 94%
Similar papers in this journal
- An updated PREDICT breast cancer prognostic model including the benefits and harms of radiotherapy 95%
- RNA Sequencing-Based Single Sample Predictors of Molecular Subtype and Risk of Recurrence for Clinical Assessment of Early-Stage Breast Cancer 94%
- Investigating the relationship between breast cancer risk factors and an AI-generated mammographic texture feature in the Nurses' Health Study II 92%
Similar papers in this journal
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 94%
- Accurate Prediction of Breast Cancer Survival through Coherent Voting Networks with Gene Expression Profiling 94%
- Accurate prognosis for localized prostate cancer through coherent voting networks and multi-omic data 93%
Similar papers in this journal
- Diffsig: Associating Risk Factors With Mutational Signatures 95%
- Causal effects of breast cancer risk factors across hormone receptor breast cancer subtypes: A two-sample Mendelian randomization study 95%
- Incorporating alternative Polygenic Risk Scores into the BOADICEA breast cancer risk prediction model 93%
Similar papers in this journal
- Heterogeneity in signaling pathway activity within primary and between primary and metastatic breast cancer 92%
- The French Early Breast Cancer Cohort (FRESH): a resource for breast cancer research and evaluations of oncology practices based on the French National Healthcare System Database (SNDS) 92%
- Genomic risk prediction for breast cancer in older women 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.