A genome-wide, machine learning-guided exploration of the cis-regulatory code involved in neuronal differentiation
Cassan, O.; Raynal, J.; Vroland, C.; Yasuzawa, K.; Kouno, T.; Chang, J.-C.; Hon, C.-C.; Shin, J. W.; Kato, M.; Takahashi, H.; Kasukawa, T.; Nobusada, T.; Lagani, V.; Lehmann, R.; Yauy, K.; Carninci, P.; Yip, C. W.; Brehelin, L.; Lecellier, C. H.
Show abstract
Gene expression is controlled by proximal and distal cis-regulatory elements (CREs), containing DNA motifs bound by various transcription factors (TFs). Other sequence features, such as specific k-mers or low complexity regions, have also been implicated [1-3]. However, in a dynamic biological process such as cell differentiation, we lack an understanding of how the transcriptional activity of CREs progressively change and what sequence features underlie these transitions, which may reflect common and/or coordinated regulatory processes. Here, we use single-cell ATAC-seq and RNA-seq to follow, at a genome scale, CREs along differentiation of induced pluripotent stem cells into cortical neurons and develop a method to automatically identify the diversity of CRE profiles and their underlying sequence features. We propose a machine-learning guided clustering algorithm, STOIC (Statistical learning TO Inform Clustering), that jointly learns an unsupervised clustering of the CREs in the space of the activity profiles and a supervised predictor associated with each cluster in the DNA-sequence space. STOIC is specifically designed to provide readily interpretable results. We show that the method identifies CRE profiles associated with highly predictive sequence features and outperforms methods solely concerned with co-activity clustering on this task. Orthogonal data collected in the same settings link the inferred CRE clusters to specific enhancer or promoter signatures. Furthermore, we show that the DNA features unveiled by STOIC reflect biologically relevant regulators and offer a valuable basis to dissect elements of the cisregulatory grammar. Finally, we demonstrate the general applicability of STOIC by analyzing five bulk CAGE datasets of human cells responding to various treatments.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 97%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 96%
- STAN, a computational framework for inferring spatially informed transcription factor activity across cellular contexts 95%
Similar papers in this journal
- Massively parallel reporter perturbation assay uncovers temporal regulatory architecture during neural differentiation 96%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 95%
Similar papers in this journal
Similar papers in this journal
- On the identification of differentially-active transcription factors from ATAC-seq data 96%
- Epigenetics is all you need: A Transformer to decode chromatin structural compartments from the epigenome 95%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.