Investigating Data Size, Sequence Diversity, and Model Complexity in MPRA-based Sequence-to-Function Prediction
Sheng, Y.; Tu, X.; Mostafavi, S.
Show abstract
We created the MPRA Dataset Collection (MDC), a curated resource of MPRA data from 12 studies comprising over 150 million labeled DNA subsequences. These datasets include both random and natural genomic sequences paired with diverse functional outputs such as gene expression and splicing efficiency. Using this collection, we build sequence-to-function (S2F) predictive models of regulatory elements and analyze these models to uncover insights into the relationship between training data requirements, experimental design, and model generalizability. By offering a high-quality, machine learning-ready repository, MDC accelerates the development of robust computational tools for deciphering the mechanisms of gene regulation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 96%
- EvoAug-TF: Extending evolution-inspired data augmentations for genomic deep learning to TensorFlow 95%
- Embeddings of genomic region sets capture rich biological associations in lower dimensions 94%
Similar papers in this journal
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
- ModiDeC: a multi-RNA modification classifier for direct nanopore sequencing 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
Similar papers in this journal
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 97%
- Differential Analysis of RNA Structure Probing Experiments at Nucleotide Resolution: Uncovering Regulatory Functions of RNA Structure 96%
- CodonTransformer: a multispecies codon optimizer using context-aware neural networks 96%
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
- Splam: a deep-learning-based splice site predictor that improves spliced alignments 96%
- Improved modeling of RNA-binding protein motifs in an interpretable neural model of RNA splicing 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.