PEAS: Detection of Clustered Differences in Genomic Data
Skola, D. D.; Glass, C.
Show abstract
MotivationStudies of genetically-distinct individuals have shown that differences in marks of transcriptional regulation such as chromatin accesibility, transcription factor binding and histone modifications are often proximally clustered along the genome. These proximal clusters, which have been labeled as cis-regulatory domains (CRDs), are thought to reflect topological features of the genome and may demarcate functional units linking genetic variation to transcriptional regulation. The problem of distinguishing CRDs from background variation is computationally difficult and current methods rely on greedy approaches with ad-hoc parameters and do not provide an assessment of statistical significance, an important consideration for investigating CRDs in small sample cohorts. ResultsWe developed a software package, PEAS (Proximal Enrichment by Approximated Sampling), to identify CRDs from a small number of samples (as few as two distinct genetic backgrounds) using a robust statistical approach. PEAS uses methods for efficient and accurate estimation of empirical distributions to quantify the significance of enriched regions, followed by a dynamic programming algorithm to identify the minimum likelihood set of non-overlapping enriched regions. We used it to identify clusters of proximally-enriched differences in the histone mark H3K27ac between two mouse strains as well as proximally-enriched regions of correlation in this mark across five mouse strains. We find that differences in histone acetylation between two mouse strains form signficant clusters that overlap closely with differences in the first principal component of their Hi-C correlation matrices. AvailabilityPEAS is written in Python and is available at https://pypi.org/project/PEAS/. Methods for approximating empirical distributions are implemented in C and Python and are available at https://pypi.org/project/empdist/.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
- Generating Correlated Data for Omics Simulation 95%
- CNAViz: An interactive webtool for user-guided segmentation of tumor DNA sequencing data 95%
Similar papers in this journal
Similar papers in this journal
- SampleQC: robust multivariate, multi-celltype, multi-sample quality control for single cell data 96%
- RoAM: computational reconstruction of ancient methylomes and identification of differentially methylated regions 96%
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 95%
Similar papers in this journal
- Variance-adjusted Mahalanobis (VAM): a fast and accurate method for cell-specific gene set scoring 95%
- S3norm: simultaneous normalization of sequencing depth and signal-to-noise ratio in epigenomic data 95%
- edgeR v4: powerful differential analysis of sequencing data with expanded functionality and improved support for small counts and larger datasets 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.