Extreme gradient boosting machine learning algorithm identifies genome-wide relevant genetic variants in prostate cancer risk prediction.
Enoma, D. O.; Osamor, V. C.; Ogunlana, O.
Show abstract
Genome-wide association studies (GWAS) identify the variants (Single Nucleotide polymorphisms) associated with a disease phenotype within populations. These genetic differences are essential in variations in incidence and mortalities, especially for Prostate cancer in the African population. Given the complexity of cancer, it is imperative to identify the variants that contribute to the development of the disease. The standard univariate analysis employed in GWAS may not capture the non-linear additive interactions between variants, which might affect the risk of developing Prostate cancer. This is because the interactions in complex diseases such as prostate cancer are usually non-linear and would benefit from a non-linear Machine Learning gradient boosting viz XGBoost (extreme gradient boosting). We applied the XGBoost algorithm and an iterative SNP selection algorithm to find the top features (SNPs) that best predict the risk of developing prostate cancer with a Support Vector Machine (SVM). The number of subjects was 907, and input features were 1,798,727 after appropriate quality control. The algorithm involved ten trials of 5-fold cross-validation to optimize the datasets hyperparameters and the prediction tasks second module (utilizing SVM). The model achieved AUC-ROC cure of 0.66, 0.57 and 0.55 on the Train, Dev and Test sets, respectively. The area under the Precision-Recall Curve was 0.69, 0.60 and 0.57 on the Train, Dev and Test sets, respectively. Furthermore, the final number of predictive risk variants was 2798, associated with 847 Ensembl genes. Interaction analysis showed that Nodes were 339 and the edges were 622 in the gene interaction network. This shows evidence that the non-linear Machine learning approach offers excellent possibilities for understanding the genetic basis of complex diseases.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 95%
- Comprehensive analysis of prostate cancer life expectancy, loss of life expectancy, and healthcare expenditures: Taiwan national cohort study spanning 2008 to 2019 94%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 94%
Similar papers in this journal
- Investigate the relevance of major signaling pathways in cancer survival using a biologically meaningful deep learning model 93%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 93%
- MCKAT, a multi-dimensional copy numbervariant kernel association test 93%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 95%
- Loss of HOXB13 expression in neuroendocrine prostate cancer 94%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.