Multi-dataset Integration and Residual Connections Improve Proteome Prediction from Transcriptomes using Deep Learning
Cranney, C. W.; Meyer, J. G.
Show abstract
Proteomes are well known to poorly correlate with transcriptomes measured from the same sample. While connected, the complex processes that impact the relationships between transcript and protein quantities remains an open research topic. Many studies have attempted to predict proteomes from transcriptomes with limited success. Here we use publicly available data from the Clinical Proteomics Tumor Analysis Consortium to show that deep learning models designed by neural architecture search (NAS) achieve improved prediction accuracy of proteome quantities from transcriptomics. We find that this benefit is largely due to including a residual connection in the architecture that allows input information to be remembered near the end of the network. Finally, we explore which groups of transcripts are functionally important for protein prediction using model interpretation with SHAP.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 95%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 95%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
Similar papers in this journal
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 95%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 93%
- driveR: A Novel Method for Prioritizing Cancer Driver Genes Using Somatic Genomics Data 93%
Similar papers in this journal
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 94%
- Optimizer's dilemma: optimization strongly influences model selection in transcriptomic prediction 94%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
- Wide and Deep Learning for Automatic Cell Type Identification 94%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.