ShapeME: A tool and web front-end for de novo discovery of structural motifs underpinning protein-DNA interactions
Schroeder, J. W.; Wolfe, M. B.; Freddolino, L.
Show abstract
Determining where transcriptional regulators bind within a genome is paramount to understanding how gene expression is regulated. Historically, position weight matrices (PWMs) have been used to define the binding preferences of DNA binding proteins1. However, PWMs treat the identity of each base in a sequence as an independent and additive measure of binding preference, which can limit their utility2. Models that consider higher order interactions between nearby bases yield greater success in predicting proteins binding to DNA, but for many proteins there is still substantial room for improvement in predicting and understanding the determinants of proteins binding to DNA3. In addition to DNA sequence motifs, structural motifs (e.g., a narrow minor groove width) are important determinants of binding for some DNA-binding proteins4. Despite the initial success of algorithms using structural features of DNA to predict binding properties of proteins from either ChIP-seq or SELEX data5-8, there remains a need for a de novo structural motif discovery framework which can be applied to data from a variety of experimental designs. Here, we present a unified workflow, capable of utilizing virtually any type of data representing sequence coverage or enrichment (e.g. ChIP-seq, RNA-seq, SELEX, etc.), to discover short structural motifs with explanatory power for a proteins DNA binding preference. We couple the DNAshapeR algorithm9 with our own information-theoretic approach to de novo motif discovery, and wrap shape and sequence motif inference and model selection into a single tool called ShapeME. Application of our structural motif discovery algorithm to proteins with ChIP-seq data in ENCODE datasets reveals a subset of proteins where short structural motifs outperform the best PWM for that protein as determined from the JASPAR database, or as identified by the sequence motif elicitation tool STREME. Our approach offers a powerful and versatile framework for inferring structural DNA binding motifs, and will complement current sequence-based motif elicitation tools in discovery of protein-DNA interaction principles. A web-based interface to ShapeME is available at https://seq2fun.dcmb.med.umich.edu/shapeme, with full source code available at https://github.com/freddolino-lab/ShapeME.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
- Inferring transcriptional regulators through integrative modeling ofpublic chromatin accessibility and ChIP-seq data 96%
- Improved modeling of RNA-binding protein motifs in an interpretable neural model of RNA splicing 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Discovering functional sequences with RELICS, an analysis method for tiling CRISPR screens 95%
- Learning And Interpreting The Gene Regulatory Grammar In A Deep Learning Framework 95%
- Global Importance Analysis: An Interpretability Method to Quantify Importance of Genomic Features in Deep Neural Networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.