GECSI: Large-scale chromatin state imputation from gene expression
Fu, J.; Ernst, J.
Show abstract
Compendiums of chromatin state annotations based on integrating maps of multiple epigenetic marks such as from ChromHMM have become a powerful resource. While these compendiums have coverage of many biological samples, there are many additional biological samples that have gene expression data but lack epigenetic mark data and chromatin state annotations. The EpiAtlas resource of the International Human Epigenome Consortium (IHEC) contains a large compendium of chromatin state annotations for which many samples have matched gene expression data, which provides the opportunity to use it to train models to predict chromatin state annotations in additional biological samples with only gene expression data available. To address this, we develop Gene Expression-based Chromatin State Imputation (GECSI), which uses a multi-class logistic regression model trained using a large compendium of gene expression and chromatin state annotations, and apply it to IHEC data. Using cross-validation, we find that GECSI accurately predicts chromatin state assignments and generates probability estimates that are predictive of observed chromatin states, overall outperforming multiple other alternative and baseline methods. GECSI-predicted chromatin states reflect relationships among biological samples and show similar transcription factor and gene annotation enrichments as observed chromatin states. Using available IHEC gene expression data, we apply GECSI to predict chromatin state annotations for 449 additional epigenomes. We expect these predicted annotations and the GECSI software will be a useful resource for chromatin state analyses in many additional biological samples.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrative epigenomic and functional characterization assay based annotation of regulatory activity across diverse human cell types 96%
- A curated benchmark of enhancer-gene interactions for evaluating enhancer-target gene prediction methods 96%
- Harmonizing single cell 3D genome data with STARK and scNucleome 96%
Similar papers in this journal
- A framework for summarizing chromatin state annotations within and identifying differential annotations across groups of samples 98%
- Learning a Pairwise Epigenomic and Transcription Factor Binding Association Score Across the Human Genome 96%
- Fast Detection of Differential Chromatin Domains with SCIDDO 95%
Similar papers in this journal
- DeepC: Predicting chromatin interactions using megabase scaled deep neural networks and transfer learning. 96%
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks 95%
- Simultaneous profiling of chromatin accessibility and methylation on human cell lines with nanopore sequencing 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.