Inferring protein from mRNA concentrations using convolutional neural networks
Schwehn, P. M.; Falter-Braun, P.
10.1101/2023.11.06.565778 bioRxivShow abstract
Transcript abundance is a widely used but poor predictor of protein abundance. As proteins are the actual agents executing biological functions, and because signaling outcome depends in a non-linear manner on the concentration of the network components, we aimed to develop a convolutional neural network-(CNN-) based predictor for Homo sapiens and the reference plant Arabidopsis thaliana. After hyperparameter optimization and initial analysis of the training data, we employed a distinct training module for value and sequence data, respectively, predicting 40% of the variance in protein levels in Homo sapiens, respectively 48% in Arabidopsis thaliana. Codon counts and peptides had the greatest predictive power. Extracting the learned weight revealed generally similar trends but also some intriguing differences between human and Arabidopsis. Many learned motifs in the 5 and 3 UTRs correspond to previously described regulatory features demonstrating that the model can learn ab initio mechanistically relevant features.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting Mean Ribosome Load for 5'UTR of any length using Deep Learning 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
- Improving deep models of protein-coding potential with a Fourier-transform architecture and machine translation task 95%
Similar papers in this journal
Similar papers in this journal
- Revisiting the Central Dogma: the distinct roles of genome, methylation, transcription, and translation on protein expression in Arabidopsis thaliana 96%
- ReorientExpress: reference-free orientation of nanopore cDNA reads with deep learning 95%
- The genetic and biochemical determinants of mRNA degradation rates in mammals 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.