IPLS-LDA: An Improved Partial Least Square Discriminant Analysis for Heterogeneous Transcriptomics and Metabolomics Data Analysis
Shahjaman, M.; Sarkar, S.; Das, S.
Show abstract
Supervised machine learning (SML) is an approach that learns from training data with known category membership to predict the unlabeled test data. There are many SML approaches in the literature and most of them use a linear score to learn its classifier. However, these approaches fail to elucidate biodiversity from heterogeneous biomedical data. Therefore, their prediction accuracies become low. Partial Least Square Linear Discriminant Analysis (PLS-LDA) is widely used in gene expression (GE) and metabolomics datasets for predicting unlabelled test data. Nevertheless, it also does not consider the non-linearity and heterogeneity pattern of the datasets. Hence, in this study, an improved PLS-LDA (IPLS-LDA) was developed by capturing the heterogeneity of datasets through an unsupervised hierarchical clustering approach. In our approach a non-linear score was calculated by combining all the linear scores obtained from the clustering method. The performance of IPLS-LDA was investigated in a comparison with six frequently used SML methods (SVM, LDA, KNN, Naive Bayes, RF, PLS-LDA) using one simulation data, one colon cancer gene expression data (GED) and one lung cancer metabolomics datasets. The resultant IPLS-LDA predictor achieved accuracy 0.841 using 10-fold cross validation in colon cancer data and accuracy 0.727 from two independent metabolomics data analysis. In both the cases IPLS-LDA outperformed other SML predictors. The proposed algorithm has been implemented in an R package, Uplsda was given in the https://github.com/snotjanu/UplsLda.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 97%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 96%
- Reconstruction of a generic genome-scale metabolic network for chicken: investigating network connectivity and finding potential biomarkers 96%
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 96%
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 96%
- Feature Extraction Approaches for Biological Sequences: A Comparative Study of Mathematical Models 96%
Similar papers in this journal
- Investigate the relevance of major signaling pathways in cancer survival using a biologically meaningful deep learning model 95%
- Shrinkage estimation of gene interaction networks in single-cell RNA sequencing data 94%
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 94%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 97%
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 96%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.