Back

RepliCNN: High-resolution inference of the DNA replication program from strand-specific 3' DNA end sequencing

Stroh, D.; Zilio, N.; Pabba, M. K.; Roukos, V.; Cardoso, M. C.; Ulrich, H. D.; Zarnack, K.

2026-03-14 bioinformatics
10.64898/2026.03.12.710907 bioRxiv
Show abstract

During S phase, the genome is replicated in a tightly regulated spatiotemporal order described as DNA replication timing (RT). Discontinuous lagging-strand synthesis produces Okazaki fragments whose strand-specific distribution reflects replication dynamics. Here, we present RepliCNN, a deep learning framework based on one-dimensional convolutional neural networks to predict RT from Okazaki fragment distributions obtained from strand-specific 3' DNA end sequencing methods such as GLOE-Seq, TrAEL-seq, or OK-Seq. RepliCNN also automatically annotates replication origins, termination zones, replication fork directionality, and origin efficiency genome-wide from a single dataset. Benchmarking on public and in-house human and yeast datasets using leave-one-chromosome-out cross-validation demonstrates high predictive accuracy in both wild-type and perturbation experiments, enabling comprehensive analyses of replication dynamics from strand-specific DNA 3' end sequencing data. HighlightsO_LIRepliCNN enables integrated analysis of replication timing, fork directionality and replication features from strand-specific 3' DNA end sequencing data. C_LIO_LIHigh-resolution replication dynamics can be inferred from a single experiment, bypassing complex multi-fraction labelling approaches. C_LIO_LIThe framework generalizes across experimental protocols, datasets, and species. C_LIO_LIThis enables cost-effective comparative analysis of replication programs across biological conditions. C_LI

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.