Back

ORCA: Predicting replication origins in circular prokaryotic chromosomes

van Meel, Z.; Baaijens, J. A.

2024-03-30 bioinformatics
10.1101/2024.03.28.587133 bioRxiv
Show abstract

The proximity of genes to the origin of replication plays a key role in replication and transcription-related processes in bacteria. Computational prediction of potential origin locations has an important role in origin discovery, critically reducing experimental costs. We present ORCA (Origin of RepliCation Assessment) as a fast and lightweight tool for the visualisation of nucleotide disparities and the prediction of the location of replication origins. ORCA uses the analysis of nucleotide disparities, dnaA-box regions, and target gene positions to find potential origin sites, and has a random forest classifier to predict which of these sites are likely origins. ORCAs prediction and visualization capabilities make it a valuable in silico method to assist in experimental determination of replication origins. ORCA is written in Python-3.11, works on any operating system with minimal effort, and can process large databases. Full implementation details are provided in the supplementary material and the source code is freely available on GitHub: https://github.com/ZoyavanMeel/ORCA.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.