PrismExp: Predicting Human Gene Function by Partitioning Massive RNA-seq Co-expression Data
Lachmann, A.; Rizzo, K.; Bartal, A.; Jeon, M.; Clarke, D. J. B.; Ma'ayan, A.
Show abstract
Gene co-expression correlations from mRNA-sequencing (RNA-seq) can be used to predict gene function based on the covariance structure that exists within such data. In the past, we showed that RNA-seq co-expression data is highly predictive of gene function and protein-protein interactions. We demonstrated that the performance of such predictions is dependent on the source of the gene expression data. Furthermore, since genes function in different cellular contexts, predictions derived from tissue-specific gene co-expression data outperform predictions derived from cross-tissue gene co-expression data. However, the identification of the optimal tissue type to maximize gene function predictions for all mammalian genes is not trivial. Here we introduce and validate an approach we term Partitioning RNA-seq data Into Segments for Massive co-EXpression-based gene function Predictions (PrismExp), for improved gene function prediction based on RNA-seq co-expression data. With coexpression data from ARCHS4, we apply PrismExp to predict a wide variety of gene functions, including pathway membership, phenotypic associations, and protein-protein interactions. PrismExp outperforms the cross-tissue co-expression correlation matrix approach on all tested domains. Hence, PrismExp can enhance machine learning methods that utilize RNA-seq coexpression correlations to impute knowledge about understudied genes and proteins.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Normalizing single-cell RNA sequencing data with internal spike-in-like genes 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
- Cluefish: mining the dark matter of transcriptional data series with over-representation analysis enhanced by aggregated biological prior knowledge 95%
Similar papers in this journal
- PRIME: a probabilistic imputation method to reduce dropouteffects in single cell RNA sequencing 95%
- HUMESS: Integrating Quantitative Transcriptomic Analysis and Metabolic Modeling to Unveil Condition-Specific Gene Signatures 95%
- Systematic analysis of alternative splicing in time course data using Spycone 95%
Similar papers in this journal
Similar papers in this journal
- SPECK: An Unsupervised Learning Approach for Cell Surface Receptor Abundance Estimation for Single Cell RNA-Sequencing Data 95%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- iTraNet: A Web-Based Platform for integrated Trans-Omics Network Visualization and Analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.