An in-depth evaluation of annotated transcription start sites in E. coli using deep learning
Clauwaert, J.; Menschaert, G.; Waegeman, W.
Show abstract
The effectiveness of deep learning methods can be largely attributed to the automated extraction of relevant features from raw data. In the field of functional genomics, this generally comprises the automatic selection of relevant nucleotide motifs from DNA sequences. To benefit from automated learning methods, new strategies are required that unveil the decision-making process of trained models. In this paper, we present several methods that can be used to gather insights on biological processes that drive any genome annotation task. This work builds upon a transformer-based neural network framework designed for prokaryotic genome annotation purposes. We find that the majority of sub-units (attention heads) of the model are specialized towards identifying DNA binding sites. Working with a neural network trained to detect transcription start sites in E. coli, we successfully characterize both locations and consensus sequences of transcription factor binding sites, including both well-known and potentially novel elements involved in the initiation of the transcription process.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 95%
- Identifying promoter sequence architectures via a chunking-based algorithm using non-negative matrix factorisation 94%
- Predicting Mean Ribosome Load for 5'UTR of any length using Deep Learning 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.