Harnessing machine learning models for epigenome to transcriptome association studies
Behjati ardakani, F.; Ashrafiyan, S.; Rumpf, L.; Hecker, D.; Schulz, M. H.
Show abstract
Understanding how epigenome variation contributes to gene expression in disease and development is a fundamental challenge. Regulatory regions show cell type-specific epigenome activity and differ in their location, size, and distance to their target genes, complicating discovery and analysis. Recent machine learning models have been proposed to address these problems by learning functions for the prediction of gene expression from epigenomic data. Here, we use the large IHEC EpiATLAS dataset to benchmark state-of-the-art linear and non-linear approaches. Each approach is optimized for over 28,000 human genes, providing a comprehensive regulatory catalog of gene models. In-depth comparison reveals that gene characteristics and the epigenomic complexity of the locus influence the difficulty of predicting the epigenome-to-transcriptome association. The model performance is further evaluated using CRISPRi and eQTL validation data. Based on these models, we conduct histone-acetylation association studies in a systematic way to investigate how epigenomic variation impacts gene expression. The model-based analysis revealed genes and regulatory regions linked to B-cell leukemia in patient data with known disease-related functions. Our work provides a foundation for applications that link epigenome variation to gene expression in human cells, by benchmarking methods on a per-gene basis, illustrating their use in a disease context and making trained models available to the community.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 96%
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 96%
- CorrAdjust unveils biologically relevant transcriptomic correlations by efficiently eliminating hidden confounders 95%
Similar papers in this journal
- WEVar: a novel statistical learning framework for predicting noncoding regulatory variants 96%
- DeepDRIM: a deep neural network to reconstruct cell-type-specific gene regulatory network using single-cell RNA-seq data 96%
- Benchmarking copy number aberrations inference tools using single-cell multi-omics datasets 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.