Cross-platform DNA motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors
Vorontsov, I. E.; Kozin, I.; Abramov, S.; Boytsov, A.; Jolma, A.; Albu, M.; Ambrosini, G.; Faltejskova, K.; Gralak, A. J.; Gryzunov, N.; Inukai, S.; Kolmykov, S.; Kravchenko, P.; Kribelbauer-Swietek, J. F.; Laverty, K. U.; Nozdrin, V.; Patel, Z. M.; Penzar, D.; Plescher, M.-L.; Pour, S. E.; Razavi, R.; Yang, A. W. H.; Yevshin, I.; Zinkevich, A.; Weirauch, M. T.; Bucher, P.; Deplancke, B.; Fornes, O.; Grau, J.; Grosse, I.; Kolpakov, F. A.; Codebook/GRECO-BIT Consortium, ; Makeev, V. J.; Hughes, T. R.; Kulakovskiy, I. V.
Show abstract
A DNA sequence pattern, or "motif", is an essential representation of DNA-binding specificity of a transcription factor (TF). Any particular motif model has potential flaws due to shortcomings of the underlying experimental data and computational motif discovery algorithm. As a part of the Codebook/GRECO-BIT initiative, here we evaluated at large scale the cross-platform recognition performance of positional weight matrices (PWMs), which remain popular motif models in many practical applications. We applied ten different DNA motif discovery tools to generate PWMs from the "Codebook" data comprised of 4,237 experiments from five different platforms profiling the DNA-binding specificity of 394 human proteins, focusing on understudied transcription factors of different structural families. For many of the proteins, there was no prior knowledge of a genuine motif. By benchmarking-supported human curation, we constructed an approved subset of experiments comprising about 30% of all experiments and 50% of tested TFs which displayed consistent motifs across platforms and replicates. We present the Codebook Motif Explorer (https://mex.autosome.org), a detailed online catalog of DNA motifs, including the top-ranked PWMs, and the underlying source and benchmarking data. We demonstrate that in the case of high-quality experimental data, most of the popular motif discovery tools detect valid motifs and generate PWMs, which perform well both on genomic and synthetic data. Yet, for each of the algorithms, there were problematic combinations of proteins and platforms, and the basic motif properties such as nucleotide composition and information content offered little help in detecting such pitfalls. By combining multiple PMWs in decision trees, we demonstrate how our setup can be readily adapted to train and test binding specificity models more complex than PWMs. Overall, our study provides a rich motif catalog as a solid baseline for advanced models and highlights the power of the multi-platform multi-tool approach for reliable mapping of DNA binding specificities. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/619379v2_ufig1.gif" ALT="Figure 1"> View larger version (61K): org.highwire.dtl.DTLVardef@79561forg.highwire.dtl.DTLVardef@54c0aorg.highwire.dtl.DTLVardef@1c33f34org.highwire.dtl.DTLVardef@16a93ba_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical AbstractC_FLOATNO C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Inferring transcriptional regulators through integrative modeling ofpublic chromatin accessibility and ChIP-seq data 96%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
Similar papers in this journal
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 95%
- Predicting gene expression from histone marks using chromatin deep learning models depends on histone mark function, regulatory distance and cellular states 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.