Soffritto: a deep-learning model for predicting high-resolution replication timing
Bolzan, D.; Ay, F.
Show abstract
MotivationReplication Timing (RT) refers to the order by which DNA loci are replicated during S phase. RT is cell-type specific and implicated in cellular processes including transcription, differentiation, and disease. RT is typically quantified genome-wide using two-fraction assays (e.g., Repli-Seq) which sort cells into early and late S phase fractions followed by DNA sequencing yielding a ratio as the RT signal. While two-fraction RT data is widely available in multiple cell lines, it is limited in its ability to capture high-resolution RT features. To address this, high-resolution Repli-Seq, which quantifies RT across 16 fractions, was developed, but it is costly and technically challenging with very limited data generated to date. ResultsHere we developed Soffritto, a deep learning model that predicts high-resolution RT data using two-fraction RT data, histone ChIP-seq data, GC content, and gene density as input. Soffritto is composed of a Long Short Term Memory (LSTM) module and a prediction module. The LSTM module learns long- and short-range interactions between genomic bins while the prediction module is composed of a fully connected layer that outputs a 16-fraction probability vector for each bin using the LSTM modules embeddings as input. By performing both within cell line and cross cell line training and testing for five human and mouse cell lines, we show that Soffritto is able to capture experimental 16-fraction RT signals with high accuracy and the predicted signals allow detection of high-resolution RT patterns. AvailabilitySoffritto is available at https://github.com/ay-lab/Soffritto.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated quality control and cell identification of droplet-based single-cell data using dropkick 95%
- Dynamic Analysis of Alternative Polyadenylation from Single-Cell RNA-Seq(scDaPars) Reveals Cell Subpopulations Invisible to Gene Expression Analysis 95%
- Comprehensive characterization of tissue-specific chromatin accessibility in L2 Caenorhabditis elegans nematodes 94%
Similar papers in this journal
- SAILER: Scalable and Accurate Invariant Representation Learning for Single-Cell ATAC-Seq Processing and Integration 95%
- A Bayesian method to cluster single-cell RNA sequencing data using Copy Number Alterations 95%
- CENTRE: A gradient boosting algorithm for Cell-type-specific ENhancer-Target pREdiction 95%
Similar papers in this journal
- Comprehensive prediction of robust synthetic lethality between paralog pairs in cancer cell lines 94%
- Computational modeling reveals cell-cycle dependent kinetics of H4K20 methylation states during Xenopus embryogenesis 94%
- Model-X knockoffs reveal data-dependent limits on regulatory network identification 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.