Harnessing DNA Foundation Models for Cross-Species Transcription Factor Binding Site Prediction in Plant Genomes
Haghani, M.; Dhulipalla, K. V.; Li, S.
Show abstract
Accurate prediction of transcription factor binding sites (TFBSs) is crucial for understanding gene regulation. While experimental methods such as ChIP-seq and DAP-seq are informative, they are labor-intensive and species-specific. Recent advancements in large-scale pretrained DNA foundation models have shown promise in overcoming these limitations. This study evaluates the performance of three such models--DNABERT-2, AgroNT, and HyenaDNA--in predicting TFBSs in plants. Using DAP-seq data from Arabidopsis thaliana and Sisymbrium irio, we benchmark their accuracy against specialized approaches, including a motif-based method and two deep learning models, DeepBind and BERT-TFBS. Our results demonstrate that foundation models, particularly HyenaDNA, offer superior predictive accuracy and computational efficiency, highlighting their potential for scalable, genome-wide TFBS prediction in plants.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 96%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 96%
- Tomtom-lite: Accelerating Tomtom enables large-scale and real-time motif similarity scoring 96%
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Biologically-relevant transfer learning improves transcription factor binding prediction 95%
- Cross-Species Prediction of Histone Modifications in Plants via Deep Learning 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.