Using The Cancer Genome Atlas from cBioPortal to Develop Genomic Datasets for Machine Learning Assisted Cancer Treatment
Asaduzzaman, A.; Thompson, C.; Sibai, F.; Uddin, M. J.
Show abstract
Predicting the impact of genetic mutations is crucial for understanding diseases like cancer. Polymorphism Phenotyping (PolyPhen) and Sorting Intolerant From Tolerant (SIFT) are key tools for assessing how amino acid substitutions affect protein function and mutation pathogenicity. To our knowledge, no ready-to-use genomic dataset exists for prediction models to identify potentially harmful mutations, which could support research and clinical decisions. This study develops genomic and non-genomic datasets using The Cancer Genome Atlas (TCGA) from cBioPortal and applies machine learning models to predict PolyPhen and SIFT scores. We explore three classification models: Random Forest (RF), Extreme Gradient Boosting (XGBoost), and an ensemble RF-XGBoost model. Experimental results show that genomic data yields more accurate predictions than non-genomic data. The ensemble RF-XGBoost model performs best on genomic data, achieving average accuracies of 88.43% for PolyPhen and 95.13% for SIFT, highlighting the potential of artificial intelligence in genetic mutation analysis for disease treatment.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multiple-kernel learning for genomic data mining and prediction 94%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 94%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 94%
Similar papers in this journal
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 94%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 94%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 93%
Similar papers in this journal
- Novel ratio-metric features enable the identification of new driver genes across cancer types 96%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 94%
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 94%
Similar papers in this journal
- MasterPATH: network analysis of functional genomics screening data 93%
- Visualizing and exploring patterns of large mutational events with SigProfilerMatrixGenerator 93%
- Genomic prediction using machine learning: A comparison of the performance of regularized regression, ensemble, instance-based and deep learning methods on synthetic and empirical data 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.