Back

Using The Cancer Genome Atlas from cBioPortal to Develop Genomic Datasets for Machine Learning Assisted Cancer Treatment

Asaduzzaman, A.; Thompson, C.; Sibai, F.; Uddin, M. J.

2025-02-22 genomics
10.1101/2025.02.17.638660 bioRxiv
Show abstract

Predicting the impact of genetic mutations is crucial for understanding diseases like cancer. Polymorphism Phenotyping (PolyPhen) and Sorting Intolerant From Tolerant (SIFT) are key tools for assessing how amino acid substitutions affect protein function and mutation pathogenicity. To our knowledge, no ready-to-use genomic dataset exists for prediction models to identify potentially harmful mutations, which could support research and clinical decisions. This study develops genomic and non-genomic datasets using The Cancer Genome Atlas (TCGA) from cBioPortal and applies machine learning models to predict PolyPhen and SIFT scores. We explore three classification models: Random Forest (RF), Extreme Gradient Boosting (XGBoost), and an ensemble RF-XGBoost model. Experimental results show that genomic data yields more accurate predictions than non-genomic data. The ensemble RF-XGBoost model performs best on genomic data, achieving average accuracies of 88.43% for PolyPhen and 95.13% for SIFT, highlighting the potential of artificial intelligence in genetic mutation analysis for disease treatment.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.