Back

Investigating Data Size, Sequence Diversity, and Model Complexity in MPRA-based Sequence-to-Function Prediction

Sheng, Y.; Tu, X.; Mostafavi, S.

2025-03-13 bioinformatics
10.1101/2025.03.11.642630 bioRxiv
Show abstract

We created the MPRA Dataset Collection (MDC), a curated resource of MPRA data from 12 studies comprising over 150 million labeled DNA subsequences. These datasets include both random and natural genomic sequences paired with diverse functional outputs such as gene expression and splicing efficiency. Using this collection, we build sequence-to-function (S2F) predictive models of regulatory elements and analyze these models to uncover insights into the relationship between training data requirements, experimental design, and model generalizability. By offering a high-quality, machine learning-ready repository, MDC accelerates the development of robust computational tools for deciphering the mechanisms of gene regulation.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.