Performance Assessment of an Unsupervised Variable Selection Approach for Biomarker Discovery and Glioblastoma Subtyping
Coletti, R.; Cerdeira, J. O.; Raydan, M.; Lopes, M. B.
Show abstract
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting task-specific variables, this approach may not be optimal when the goal is to preserve the dataset structure and the overall biological information for multiple downstream analyses. In such cases, a priori unsupervised variable selection is preferable. In this study, we evaluate several unsupervised variable selection approaches to derive a representative subset of the original dataset. Building on the performance assessment results, we introduce TRIM-IT, a novel tool for unsupervised variable selection, clustering, survival analysis, and differential gene expression analysis. Applied to glioblastoma (GBM) data, TRIM-IT identified three clusters that correlate with tumor histology, exhibit distinct survival curves, and display unique molecular profiles with genes potentially serving as biomarkers. The tool is available for reproduction and adaptation to other studies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 96%
- A Novel Tool for Multi-Omics Network Integration and Visualization: A Study of Glioma Heterogeneity 95%
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 95%
Similar papers in this journal
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 97%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 94%
- Identification of high-risk COVID-19 patients using machine learning 94%
Similar papers in this journal
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 95%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 94%
- Identification of miRNA signatures for kidney renal clear cell carcinoma using the tensor-decomposition method 94%
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 96%
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 95%
- Assessing Random Forest self-reproducibility for optimal short biomarker signature discovery 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.