Model-based Standardization of Correlation Coefficients Improves Multi-Omic Clustering and Biological Signal Discovery
Robinson, M.; Noh, H.; Pflieger, L.; Rappoport, N.
Show abstract
Multi-omic data pose a particular challenge for Weighted Correlation Network Analysis (WCNA or WGCNA) due to (platform- or) batch-specific characteristics, such as resolution, accuracy, dynamic range, and sources of spurious variation. When unaccounted for, these differences can result in a bias toward single-batch clusters as well as greater sensitivity to "noisier" batches during clustering. Here we propose mitigating these effects using null models fitted separately to the bulk of analyte-analyte correlations within each batch and across each pair of batches. We then map the batch-specific null models to a standard null model, removing batch-dependent distributional differences. This approach is compatible with any correlation-based clustering approach. Since the null model represents information not captured in individual pairwise correlations, we show how to incorporate this additional information into both distance-based clustering and WCNA. For distance-based clustering, we increase distances corresponding to correlations consistent with the null model. For WCNA, we provide a new soft threshold (adjacency) function based on the likelihood of a correlation under the null model. The resulting network can be easily incorporated into the WCNA workflow. These methods are implemented in R package standardcor, and we illustrate the package on simulated data as well as an existing multi-omic dataset.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Optimal construction of a functional interaction network from pooled library CRISPR fitness screens 94%
- Single sample pathway analysis in metabolomics: performance evaluation and application 94%
- Guidelines for cell-type heterogeneity quantification based on a comparative analysis of reference-free DNA methylation deconvolution software 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.