Inferring binding specificities of human transcription factors with the wisdom of crowds
Gryzunov, N.; Penzar, D.; Kamenets, V.; Vyaltsev, V.; Kozin, I.; Eliseeva, I. A.; Nozdrin, V.; Vorontsov, I. E.; Bushuev, S.; Strekalovskikh, V.; Zinkevich, A.; Andrews, G.; Bedarew, M.; Blass, I.; Frolov, D.; Lariushina, I.; Moore, J.; Orenstein, Y.; Roev, G.; Salimov, D.; Shimshoviz, N.; Tziony, I.; Weng, Z.; IBIS Consortium, ; GRECO-BIT/Codebook Consortium, ; Bucher, P.; Deplancke, B.; Fornes, O.; Grau, J.; Grosse, I.; Jolma, A.; Kolpakov, F. A.; Makeev, V. J.; Hughes, T. R.; Kulakovskiy, I. V.
Show abstract
DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=175 SRC="FIGDIR/small/688692v1_ufig1.gif" ALT="Figure 1"> View larger version (64K): org.highwire.dtl.DTLVardef@1c6677corg.highwire.dtl.DTLVardef@b4124aorg.highwire.dtl.DTLVardef@1ce2b1org.highwire.dtl.DTLVardef@66e917_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep DNAshape: Predicting DNA shape considering extended flanking regions using a deep learning method 95%
- Massively parallel reporter perturbation assay uncovers temporal regulatory architecture during neural differentiation 95%
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 95%
Similar papers in this journal
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Correcting gradient-based interpretations of deep neural networks for genomics 95%
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- Predicting gene expression from histone marks using chromatin deep learning models depends on histone mark function, regulatory distance and cellular states 95%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.