Back

Using Deep Learning to predict replication timing reveals baseline control of genomic DNA sequence

Janakievski, N.; Joubert, P. M.; Poetsch, A. R.

2026-07-22 genomics
10.64898/2026.07.18.739304 bioRxiv
Show abstract

Human DNA replicates according to a precise schedule: certain regions are replicated early, while others replicate later in the cell cycle. The DNA sequence and the epigenome are both indicative for replication timing (RT), their relative contributions and potential complementarity yet remain unclear. Here, we utilize Repli-Seq data and machine learning to investigate these relationships in seven human cell lines. We show that GROVER (Genome Rules Obtained Via Extracted Representations), a DNA language model, trained purely on DNA sequence can be fine-tuned for RT, which indicates that information on DNA sequence is largely sufficient to predict RT. Integrating GROVERs sequence predictions as an additional feature with diverse epigenetic sequencing data, yields a multimodal model that outperforms approaches that use exclusively the DNA sequence or the epigenome, which demonstrates predictive synergy. By evaluating input feature contributions, we find that epigenetic features vary in their RT prediction contribution. Some epigenetic features see their contribution diminish in the presence of DNA sequence information, others remain predictive. GROVERs sequence representations emerge as the most dominant feature. Furthermore, to assess cell-type specificity of RT, we partitioned the genome into constitutive domains that replicate uniformly across cell lines and cell-type-specific domains that shift during differentiation. While DNA sequence features contribute more to constitutive predictions, they also remain informative within cell-type-specific regions. Together, our findings dissect the relationship between DNA sequence and the epigenome relative to RT and suggest that primary DNA sequence instructions exert a powerful baseline control over constitutive and cell-type-specific RT, fine-tuned by regulatory cues beyond the DNA sequence.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.