Multinomial Convolutions for Joint Modeling of Sequence Motifs and Enhancer Activities
Park, M.; Singh, S.; Grisanti Canozo, F. J.; Samee, M. A. H.
Show abstract
Massively parallel reporter assays (MPRAs) have enabled the study of transcriptional regulatory mechanisms at an unprecedented scale and with high quantitative resolution. However, this realm lacks models that can discover sequence-specific signals de novo from the data and integrate them in a mechanistic way. We present MuSeAM (Multinomial CNNs for Sequence Activity Modeling), a convolutional neural network that overcomes this gap. MuSeAM utilizes multinomial convolutions that directly model sequence-specific motifs of protein-DNA binding. We demonstrate that MuSeAM fits MPRA data with high accuracy and generalizes over other tasks such as predicting chromatin accessibility and prioritizing potentially functional variants.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 98%
- Biologically-relevant transfer learning improves transcription factor binding prediction 97%
- A base-resolution panorama of the in vivo impact of cytosine methylation on transcription factor binding 97%
Similar papers in this journal
- On the identification of differentially-active transcription factors from ATAC-seq data 96%
- Epigenetics is all you need: A Transformer to decode chromatin structural compartments from the epigenome 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 96%
Similar papers in this journal
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 97%
- The adapted Activity-By-Contact model for enhancer-gene assignment and its application to single-cell data 97%
- Tomtom-lite: Accelerating Tomtom enables large-scale and real-time motif similarity scoring 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.