High-dimensional association detection in large scale genomic data
Koch, H.; Keller, C. A.; Giardine, B.; Xiang, G.; Zhang, F.; Wang, Y.; Hardison, R. C.; Li, Q.
Show abstract
Joint analyses of genomic datasets obtained in multiple different conditions are essential for understanding the biological mechanism that drives tissue-specificity and cell differentiation, but they still remain computationally challenging. To address this we introduce CLIMB (Composite LIkelihood eMpirical Bayes), a statistical methodology that learns patterns of condition-specificity present in genomic data. CLIMB provides a generic framework facilitating a host of analyses, such as clustering genomic features sharing similar condition-specific patterns and identifying which of these features are involved in cell fate commitment. We apply CLIMB to three sets of hematopoietic data, which examine CTCF ChIP-seq measured in 17 different cell populations, RNA-seq measured across constituent cell populations in three committed lineages, and DNase-seq in 38 cell populations. Our results show that CLIMB improves upon existing alternatives in statistical precision, while capturing interpretable and biologically relevant clusters in the data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A statistical framework for differential pseudotime analysis with multiple single-cell RNA-seq samples 96%
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 96%
- A robust model for cell type-specific interindividual variation in single-cell RNA sequencing data 96%
Similar papers in this journal
- scBFA: modeling detection patterns to mitigate technical noise in large-scale single cell genomics data 96%
- PreTSA: computationally efficient modeling of temporal and spatial gene expression patterns 96%
- GoM DE: interpreting structure in sequence count data with differential expression analysis allowing for grades of membership 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.