Back

Inferring binding specificities of human transcription factors with the wisdom of crowds

Gryzunov, N.; Penzar, D.; Kamenets, V.; Vyaltsev, V.; Kozin, I.; Eliseeva, I. A.; Nozdrin, V.; Vorontsov, I. E.; Bushuev, S.; Strekalovskikh, V.; Zinkevich, A.; Andrews, G.; Bedarew, M.; Blass, I.; Frolov, D.; Lariushina, I.; Moore, J.; Orenstein, Y.; Roev, G.; Salimov, D.; Shimshoviz, N.; Tziony, I.; Weng, Z.; IBIS Consortium, ; GRECO-BIT/Codebook Consortium, ; Bucher, P.; Deplancke, B.; Fornes, O.; Grau, J.; Grosse, I.; Jolma, A.; Kolpakov, F. A.; Makeev, V. J.; Hughes, T. R.; Kulakovskiy, I. V.

2025-11-17 bioinformatics
10.1101/2025.11.16.688692 bioRxiv
Show abstract

DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=175 SRC="FIGDIR/small/688692v1_ufig1.gif" ALT="Figure 1"> View larger version (64K): org.highwire.dtl.DTLVardef@1c6677corg.highwire.dtl.DTLVardef@b4124aorg.highwire.dtl.DTLVardef@1ce2b1org.highwire.dtl.DTLVardef@66e917_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.