ORCA: Predicting replication origins in circular prokaryotic chromosomes
van Meel, Z.; Baaijens, J. A.
Show abstract
The proximity of genes to the origin of replication plays a key role in replication and transcription-related processes in bacteria. Computational prediction of potential origin locations has an important role in origin discovery, critically reducing experimental costs. We present ORCA (Origin of RepliCation Assessment) as a fast and lightweight tool for the visualisation of nucleotide disparities and the prediction of the location of replication origins. ORCA uses the analysis of nucleotide disparities, dnaA-box regions, and target gene positions to find potential origin sites, and has a random forest classifier to predict which of these sites are likely origins. ORCAs prediction and visualization capabilities make it a valuable in silico method to assist in experimental determination of replication origins. ORCA is written in Python-3.11, works on any operating system with minimal effort, and can process large databases. Full implementation details are provided in the supplementary material and the source code is freely available on GitHub: https://github.com/ZoyavanMeel/ORCA.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 95%
- PIPETS: A statistically informed, gene-annotation agnostic analysis method to study bacterial termination using 3'-end sequencing. 95%
- ARYANA-BS: Context-Aware Alignment of Bisulfite-Sequencing Reads 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.