Sequence based prediction of cell type specific microRNA binding and mRNA degradation for therapeutic discovery
Kanuparthi, B.; Pour, S. E.; Findlay, S. D.; Wagih, O.; Gutierrez, J. M.; Gao, R.; Wintersinger, J.; Lin, J.; Gabra, M.; Bohn, E.; Lau, T.; Cole, C. B.; Jung, A.; Celaj, A.; Soares, F.; Gray, R.; Vaz, B.; Delfosse, K.; Lodaya, V.; Bhargava, S.; Ly, D.; Yusuf, F.; Kron, K.; Hoffman, G.; Gandhi, S.; Frey, B. J.
Show abstract
MicroRNAs and RNA binding proteins are crucial elements of post-transcriptional gene regulation, which governs the fate of mRNA molecules in the cell. However, the landscape of these regulatory interactions, particularly across different mammalian cell types, remains underexplored. We describe REPRESS, a deep learning model that predicts cell-type-specific microRNA binding and mRNA degradation directly from RNA sequence. REPRESS was trained on AGO2-CLIP, miR-eCLIP and Degradome-Seq data profiling millions of microRNA binding and mRNA degradation sites across multiple cell types in human and mouse. It reveals biology that other state-of-the-art methods did not, such as identifying repressive non-canonical miRNA target sites and decoding the regulatory effects of sequence context and miRNA binding site multiplicity. REPRESS outperforms other advanced methods and neural architectures on a comprehensive suite of seven orthogonal tasks, including identifying genetic variants that affect microRNA binding, predicting out-of-distribution data from massively parallel reporter assays, and predicting canonical and non-canonical miRNA mediated repression. To demonstrate the general utility of REPRESS, we show that it provides insights into novel biology and the design of RNA therapeutics. Code is available at : https://github.com/deepgenomics/repress
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The genetic and biochemical determinants of mRNA degradation rates in mammals 97%
- Towards In-Silico CLIP-seq: Predicting Protein-RNA Interaction via Sequence-to-Signal Learning 97%
- ZetaSuite, A Computational Method for Analyzing Multi-dimensional High-throughput Data, Reveals Genes with Opposite Roles in Cancer Dependency 96%
Similar papers in this journal
- Uncalled4 improves nanopore DNA and RNA modification detection via fast and accurate signal alignment 96%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 95%
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks 95%
Similar papers in this journal
- A high-resolution map of functional miR-181 response elements in the thymus reveals the role of coding sequence targeting and an alternative seed match 96%
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
- Using single-cell perturbation screens to decode the regulatory architecture of splicing factor programs 95%
Similar papers in this journal
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- Optimizing 5'UTRs for mRNA-delivered gene editing using deep learning 95%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.