Transfer learning and DNA language models enhance transcription factor binding predictions
Aksu, E. D.; Vingron, M.
Show abstract
Identification of in vivo transcription factor (TF) binding sites is crucial to understand gene regulatory networks, but the lack of scalability in the methods for their experimental identification directs researchers towards computational models. TF binding site prediction models are often specific for a given TF, which also hinders the generalizability of models to previously unseen TFs. Here, we present an approach to predict in vivo TF binding sites using DNA accessibility, TF RNA expression and TF binding motifs. Our novel method leverages DNA language model embeddings and transfer learning to improve its accuracy and generalizability, achieving a mean area under the precision-recall curve (AUPR) of 0.51 in held-out cell types and chromosomes in the ENCODE-DREAM in vivo TFBS prediction challenge, outperforming the top-ranked methods. Furthermore, we show that prediction accuracy increases when TFs are highly active and exhibit cell-type specific expression. We finally test our models in an independent dataset on previously unseen TFs, and report a mean AUPR of 0.36, which is state-of-the-art in a cross-TF, cross-cell type and cross-chromosomal setting.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Biologically-relevant transfer learning improves transcription factor binding prediction 97%
- Inferring transcriptional regulators through integrative modeling ofpublic chromatin accessibility and ChIP-seq data 97%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 97%
Similar papers in this journal
- Massively parallel reporter perturbation assay uncovers temporal regulatory architecture during neural differentiation 97%
- Beyond accessibility: ATAC-seq footprinting unravels kinetics of transcription factor binding during zygotic genome activation 96%
- A supervised learning framework for chromatin loop detection in genome-wide contact maps 96%
Similar papers in this journal
- Predicting gene expression from histone marks using chromatin deep learning models depends on histone mark function, regulatory distance and cellular states 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- Leveraging three-dimensional chromatin architecture for effective reconstruction of enhancer-target gene regulatory network 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.