Machine Learning Models for Predicting Multiple Myeloma Staging and MGUS Progression Using Gene Expression Data
Karathanasis, N.; Spyrou, G. M.
Show abstract
In this study, we developed and evaluated Machine Learning (ML) models aimed at predicting the stage of multiple myeloma (MM) and the progression of monoclonal gammopathy of undetermined significance (MGUS) to MM. Accurate staging of MM is critical for determining appropriate treatment strategies, and our models, employing algorithms such as ElasticNet, Random Forest, Boosting, and Support Vector Machines, demonstrated high efficacy in capturing the biological differences across disease stages. Among these, the ElasticNet model exhibited strong generalizability, achieving consistent multiclass AUC values across various datasets and data transformations. Predicting MGUS progression to MM presents a significant challenge due to the scarcity of MGUS cases that have progressed. We employed a two-pronged approach to address this: developing models using a limited dataset containing progressing MGUS patients and training models on combined MGUS and MM datasets. The models achieved AUC values slightly above 0.8, particularly with ElasticNet, Boosting and Support Vector Machines, indicating their potential in stratifying MGUS patients by progression risk. This study is original in integrating MM data with MGUS cases to enhance the predictive accuracy of MGUS progression, offering a novel methodology with potential clinical applications in patient monitoring and early intervention. Our feature selection and enrichment analyses further revealed that the identified genes are involved in key signaling pathways, including PI3K-Akt, MAPK, Wnt, and mTOR, all of which play crucial roles in MM pathogenesis. These findings align with established biological knowledge, suggest possible therapeutic targets and increase the explainability of our models.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting Prognosis and IDH Mutation Status for Patients with Lower-Grade Gliomas Using Whole Slide Images 93%
- Survival Genie, a web platform for survival analysis across pediatric and adult cancers. 92%
- Genome-wide investigation of gene-cancer associations for the prediction of novel therapeutic targets in oncology 92%
Similar papers in this journal
- A Comprehensive Targeted Panel of 295 Genes: Unveiling Key Disease Initiating and Transformative Biomarkers in MultipleMyeloma 96%
- Computational flow cytometry immunophenotyping at diagnosis is unable to predict relapse in childhood B-cell Acute Lymphoblastic Leukemia 94%
- Whole slide image representation in bone marrow cytology 92%
Similar papers in this journal
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 92%
- Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomics 92%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 92%
Similar papers in this journal
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 92%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 91%
- driveR: A Novel Method for Prioritizing Cancer Driver Genes Using Somatic Genomics Data 91%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 93%
- Wide and Deep Learning for Automatic Cell Type Identification 92%
- Explainable Machine Learning for Preoperative Relapse Prediction in Molecularly Stratified Endometrial Cancer: A Single-Center Finnish Cohort Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.