Polygenic Risk Score for Gastric Cancer
Pratinidhi, S. P.; Sayad, S.
Show abstract
BackgroundGastric Cancer is one of the most predominant types of cancer in the world, and its genomic links are currently being studied at great depth. In this paper, we work towards using Genome Wide Association Studies (GWAS) data for identifying the Single Nucleotide Polymorphisms (SNPs) which have the strongest correlation with the occurrence of gastric cancer through statistical tests and to leverage them to build a predictive model using machine learning algorithms. Polygenic risk scoring (PRS) is a straightforward predictive model for assigning genetic risk to individual outcomes (cancer or healthy). MethodGenome Wide Association Studies (GWAS) data for Gastric Cancer was subjected to different statistical tests. Chi-square was used for feature selection by determining the degree of association between each probe (SNP) and the target (cancer or control). These results were used to eliminate many probes and proceed with only those that are statistically significant. Naive Bayes Classifier and Catboost machine learning algorithms were used to build classification models to predict (score) gastric cancer. ResultsNaive Bayes classifier and Catboost classification algorithms were used for modeling. The features were selected by performing Chi-square test on each of the 319283 SNPs in the data. These values were then ordered according to the negative log of the p-value and the top 5, 100 and 1000 features were used as inputs in the classification models. The Naive Bayes classifier gave an accuracy in the range of 0.60 to 0.76 for different sets of features. The Catboost algorithm proved to be more suited for this application as it gave an accuracy above 0.90 for all subsets of features. ConclusionsThis paper aims at creating a highly accurate classification model to predict the occurrence of gastric cancer from GWAS genome data. The Catboost model with an input space of 100 SNPs yielded the best results with an accuracy of 0.93 and can be considered as a polygenic risk scoring model to score new patients for gastric cancer.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles. 94%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 94%
- Expression based biomarkers and models to classify early and late stage samples of Papillary Thyroid Carcinoma 93%
Similar papers in this journal
- Genetic Risk Factors for Colorectal Cancer in Multiethnic Indonesians 95%
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 95%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 93%
Similar papers in this journal
- PanClassif: Improving pan cancer classification of single cell RNA-seq using machine learning 94%
- Gene expression profiles and pathway enrichment analysis to identification of differentially expressed gene and signaling pathways in epithelial ovarian cancer based on high-throughput RNA-seq data 93%
- Comparative Analysis of Human Coronaviruses Focusing on Nucleotide Variability and Synonymous Codon Usage Pattern 91%
Similar papers in this journal
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 93%
- Viral miRNAs Confer Survival in Host Cells by Targeting Apoptosis Related Host Genes 91%
- Finding Consensus miRNAs Silencing KLF1 Expression as A Promising Therapeutic Option of Sickle Cell Anemia 91%
Similar papers in this journal
- The Impact of Fasting the Holy Month of Ramadan on Colorectal Cancer Patients and Two Tumor Biomarkers: A Tertiary-Care Hospital Experience 93%
- Impact of Losartan on Portal hypertension and Liver Cirrhosis: A Systematic Review 90%
- Comparative study between first and second wave of COVID-19 deaths in India - a single center study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.