TFClassPredict: A Novel Deep Learning Framework for Transcription Factor Binding Site Analysis Using Evolutionarily Conserved DNA-Binding Domain Annotations
Timucin, C. H.; Ickes, C.; Conrads, K.; Harms, B. C.; Beissbarth, T.; Haubrock, M.
Show abstract
Transcription factors (TFs) are proteins that regulate gene expression by binding to short specific sequences in DNA. The binding of TFs to DNA is actualized through their DNA-binding domains (DBDs). The interactions between TFs and DNA are fundamental for understanding gene regulation mechanisms, which form the basis of many cellular activities and processes. While many models have been developed to predict TF binding, there is a lack of a comprehensive model that accounts for the similarity of binding characteristics of TFs within the same DBD families. Our model, TFClassPredict, introduces a novel approach to identify transcription factor binding sites (TFBSs) based on the structural annotations of evolutionarily conserved DNA-binding domains (DBDs). By leveraging strong canonical binding patterns, TFClassPredict provides high-confidence predictions that are crucial for reliable regulatory analysis. By fine-tuning the DNABERT model, TFClassPredict was developed to classify DNA sequences across different hierarchical levels, with 6 Superclasses at the top level and 17 more specific Class-level annotations, reflecting different degrees of DBD similarity. TFClassPredict achieved high performance across both hierarchical levels, with an average AUC of 0.988, precision of 0.942, and recall of 0.906 at the SuperClass level, and an average AUC of 0.990, precision of 0.937, and recall of 0.914 at the Class-level. TFClassPredict demonstrated its ability to reveal distinct regulatory landscapes associated with cancer progression. The Class-level model is publicly available for use and can be accessed at https://gitlab.gwdg.de/hti/tfclass_dnabert.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data 95%
- Mixing features of transcription factors and genes enables accurate prediction of gene regulation relationships for unknown transcription factors 94%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 94%
Similar papers in this journal
- Dissecting the binding mechanisms of transcription factors to DNA using a statistical thermodynamics framework. 95%
- LBFextract: unveiling transcription factor dynamics from liquid biopsy data 95%
- Modeling and analysis of site-specific mutations in cancer identifies known plus putative novel hotspots and bias due to contextual sequences 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.