CREsted: modeling genomic and synthetic cell type-specific enhancers across tissues and species
Kempynck, N.; De Winter, S.; Blaauw, C. H.; Konstantakos, V.; Dieltiens, S.; Eksi, E. C.; Bercier, V.; Taskiran, I. I.; Hulselmans, G.; Spanier, K.; Christiaens, V.; Van Den Bosch, L.; Mahieu, L.; Aerts, S.
Show abstract
Sequence-based deep learning models have become the state of the art for the analysis of the genomic regulatory code. Particularly for transcriptional enhancers, deep learning models excel at deciphering sequence features and grammar that underlie their spatiotemporal activity. To enable end-to-end enhancer modeling and design, we developed a software and modeling package, called CREsted. It combines preprocessing starting from single-cell ATAC-seq data; modeling with a choice of several architectures for training classification and regression models on either topics or pseudobulk peak heights; sequence design using multiple strategies; and downstream analysis through a collection of tools to locate transcription factor (TF) binding sites, infer the effect of a TF (activating or repressing) on enhancer accessibility, decipher enhancer grammar, and score gene loci. We demonstrate CREsted using a mouse cortex model that we validate using the BICCN collection of in vivo validated mouse brain enhancers. Classical enhancers in immune cells, including the IFNB1 enhanceosome are revisited using a PBMC model, and we assess the accuracy of TF binding site predictions with ChIP-seq. Additionally, we use CREsted to compare mesenchymal-like cancer cell states between tumor types; and we investigate different fine-tuning strategies of Borzoi within CREsted, comparing their performance and explainability with CREsted models trained from scratch. Finally, we train a CREsted model on a scATAC-seq atlas of zebrafish development and use this to design and in vivo validate cell type-specific synthetic enhancers in three tissues. For varying datasets, we demonstrate that CREsted facilitates efficient training and analyses, enabling scrutinization of the enhancer logic and design of synthetic enhancers across tissues and species. CREsted is available at https://crested.readthedocs.io.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sequence-based modeling of genome 3D architecture from kilobase to chromosome-scale 98%
- Integrated annotation and analysis of genomic features reveal new types of functional elements and large-scale epigenetic phenomena in the developing zebrafish 98%
- Dynamic network-guided CRISPRi screen reveals CTCF loop-constrained nonlinear enhancer-gene regulatory activity in cell state transitions 97%
Similar papers in this journal
- CREaTor: zero-shot cis-regulatory pattern modeling with attention mechanisms 97%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 97%
Similar papers in this journal
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks 98%
- DeepC: Predicting chromatin interactions using megabase scaled deep neural networks and transfer learning. 97%
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 97%
Similar papers in this journal
- Massively parallel characterization of transcriptional regulatory elements in three diverse human cell types 98%
- AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model 98%
- A human DNA methylation atlas reveals principles of cell type-specific methylation and identifies thousands of cell type-specific regulatory elements 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.