Multiblock LASSO Framework for Cancer Gene Selection from RNA-Seq PANCAN Data
Ashraf, Z. A.; Aslam, M.; Mehmood, T.; Al-Essa, L. A. A.
Show abstract
The Cancer RNA-HiSeq PANCAN dataset consists of RNA-Seq gene expression data collected from multiple cancer types. It is a high-dimensional dataset, meaning it has thousands of gene expression features (predictors) and relatively fewer samples (observations). The dataset contains thousands of genes, making it difficult to identify key biomarkers. In order to reduce data and comprehend the modeled link, variable selection is essential. Least Regression using Absolute Shrinkage and Selection Operator (LASSO) is one modeling technique that deals with high throughput data. The data might be divided into different blocks representing different biological pathways or cancer types. Many genes are correlated, which can reduce interpretability. In many areas, including modern biology, variable selection is an important problem. For instance, choosing genetic characteristics for categorization (i.e., identifying harmful bacteria, diagnosing diseases, etc.) is an example of this. Multiblock Lasso (a variant of Lasso regression) is particularly useful when data is structured into blocks (e.g., different biological processes or pathways). It helps in selecting important features across multiple blocks, improving interpretability by grouping related genes, reducing over fitting in high-dimensional datasets. In this study, we apply Multiblock Lasso to extract significant gene features for cancer classification. We preprocess the dataset, define block structures using biological pathways, and optimize the regularization parameters using cross-validation. Experimental results demonstrate that Multiblock Lasso effectively reduces dimensionality while maintaining classification accuracy, making it a powerful tool for biomarker discovery in cancer genomics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 96%
- An R Package for Divergence Analysis of Omics Data 96%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 97%
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 96%
- High Dimensionality Reduction by Matrix Factorization for Systems Pharmacology 96%
Similar papers in this journal
- Stochastic LASSO for extremely high-dimensional genomic data 96%
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 95%
- Robust Feature Selection for Cancer Microarray Data Using a Hybrid mRMR and Binary Lion Optimization Algorithm 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.