snATAC-Express infers Gene Expression from Prioritized Chromatin Accessibility Peaks using Machine Learning
Brown, M.; Ferrari, A.; Dodd, A.; Shi, F.; Kolachala, V. L.; Kugathasan, S.; Wolfinger, R. D.; Gibson, G.
Show abstract
BackgroundSingle cell multi-omic investigation opens-up new opportunities to understand mechanisms of gene regulation. Existing methods for inferring transcript abundance from chromatin accessibility fail to prioritize the most relevant peaks and tend to assume positive associations between ATAC peaks and RNA counts. We hypothesize that gene regulation can be modeled as a function of combined positive and negative interactions among peaks and that causal regulatory variants are enriched in the vicinity of the most critical peaks. ResultsA machine learning pipeline leveraging single nuclear multiomic transcriptome and chromatin accessibility data is developed to model gene expression as a function of ATAC peak intensity. Multiome data was available for 18 immune cell types from 29 donors, 19 with Crohns disease. The pipeline aggregates results from three machine learning approaches (random forest regression, XGBoost, and Light GBM) as well as linear regression to identify which ATAC peaks contribute to explaining variation among donors and cell types in pseudobulk gene expression. The coefficient of determination with cross-validation was used to identify robust models which typically explain between 5% and 40% of transcript abundance, utilizing on average 47% of the ATAC peaks, representing a significant gain in predictive accuracy. The most important peaks are enriched in GWAS variants for inflammatory bowel disease and the autoimmune disease systemic lupus erythematosus, but not for rheumatoid arthritis. ConclusionAtlanta Plots visualize the proportion of ATAC peaks contributing to a predictive model of gene expression as well as the proportion of variance explained by the model. Software implementing our pipeline, "snATAC-Express", is freely available on GitHub.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COCOA: Coordinate covariation analysis of epigenetic heterogeneity 96%
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 95%
- Allele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs 95%
Similar papers in this journal
Similar papers in this journal
- Characterization of a strain-specific CD-1 reference genome reveals potential inter- and intra-strain functional variability 95%
- Neural network modeling of differential binding between wild-type and mutant CTCF reveals putative binding preferences for zinc fingers 1-2 94%
- Copy number normalization distinguishes differential signals driven by copy number differences in ATAC-seq and ChIP-seq 94%
Similar papers in this journal
Similar papers in this journal
- MUFFIN : A suite of tools for the analysis of functional sequencing data 95%
- Towards Personalized Epigenomics: Learning Shared Chromatin Landscapes and Joint De-Noising of Histone Modification Assays 94%
- Identifying similar populations across independent single cell studies without data integration 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.