Uncovering Cross-Cohort Molecular Features with Multi-Omics Integration Analysis
Jiang, M.-Z.; Aguet, F.; Ardlie, K.; Chen, J.; Cornell, E.; Cruz, D.; Durda, P.; Gabriel, S. B.; Gerszten, R. E.; Guo, X.; Johnson, C. W.; Kasela, S.; Lange, L. A.; Lappalainen, T.; Liu, Y.; Reiner, A. P.; Smith, J.; Sofer, T.; Taylor, K. D.; Tracy, R. P.; VanDenBerg, D. J.; Wilson, J. G.; Rich, S. S.; Rotter, J. I.; Love, M. I.; Raffield, L. M.; Li, Y.
Show abstract
Integrative approaches that simultaneously model multi-omics data have gained increasing popularity because they provide holistic system biology views of multiple or all components in a biological system of interest. Canonical correlation analysis (CCA) is a correlation-based integrative method. It was initially designed to extract latent features shared between two assays by finding the linear combinations of features - referred to as canonical vectors (CVs) - within each assay that achieve maximal across-assay correlation. Sparse multiple CCA (SMCCA), a widely-used derivative of CCA, allows more than two assays but can result in non-orthogonal CVs when applied to high-dimensional data. Here, we incorporated a variation of the Gram-Schmidt (GS) algorithm with SMCCA to improve orthogonality among CVs. Applying our SMCCA-GS method to proteomics and methylomics data from the Multi-Ethnic Study of Atherosclerosis (MESA) and Jackson Heart Study (JHS), we identified strong associations between blood cell counts and protein abundance. This finding suggests that adjustment of blood cell composition should be considered in protein-based association studies. Importantly, CVs obtained from two independent cohorts demonstrate transferability across the cohorts. For example, proteomic CVs learned from JHS explain similar amounts of blood cell count phenotypic variance in MESA, explaining 39.0% ~ 50.0% variation in JHS and 38.9% ~ 49.1% in MESA, similar transferability was observed for other omics-CV-trait pairs. This suggests that biologically meaningful and cohort-agnostic variation is captured by CVs. We further developed Sparse Supervised Multiple CCA (SSMCCA) to allow supervised integration analysis for more than two assays. We anticipate that applying our SMCCA-GS and SSMCCA on various cohorts would help identify cohort-agnostic biologically meaningful relationships between multi-omics data and phenotypic traits. Author SummaryComprehensive understanding of human complex traits may benefit from incorporation of molecular features from multiple biological layers such as genome, epigenome, transcriptome, proteome, and metabolome. CCA is a correlation-based method for multi-omics data which reduces the dimension of each omic assay to several orthogonal components - commonly referred to as canonical vectors (CVs). The widely-used SMCCA method allows effective dimension reduction and integration of multi-omics data, but suffers from potentially highly correlated CVs when applied to high-dimensional omics data. Here, we improve the statistical independence among the CVs by adopting a variation of the GS algorithm. We applied our SMCCA-GS method to proteomic and methylomic data from two cohort studies, MESA and JHS. Our results reveal a pronounced effect of blood cell counts on protein abundance, strongly suggesting blood cell composition adjustment in protein-based association studies may be necessary. Finally, we present SSMCCA which allows supervised CCA analysis for the association between one phenotype of interest and more than two assays. We anticipate that SMCCA-GS would help reveal meaningful system-level factors from biological processes involving features from multiple assays; and SSMCCA would further empower interrogation of these factors for phenotypic traits related to health and diseases.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RADAR: Differential analysis of MeRIP-seq data with a random effect model 95%
- Missing cell types in single-cell references impact deconvolution of bulk data but are detectable 94%
- A systematic evaluation of 41 DNA methylation predictors across 101 data preprocessing and normalization strategies highlights considerable variation in algorithm performance 94%
Similar papers in this journal
- SSMD: A semi-supervised approach for a robust cell type identification and deconvolution of mouse transcriptomics data 94%
- BayeSMART: Bayesian Clustering of Multi-sample Spatially Resolved Transcriptomics Data 94%
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 94%
Similar papers in this journal
- Nonlinear ridge regression improves cell-type-specific differential expression analysis 94%
- Sensei: How many samples to tell evolution in single-cell studies? 94%
- Decoding Single-Cell Multiomics: scMaui - A Deep Learning Framework for Uncovering Cellular Heterogeneity in Presence of Batch Effects and Missing Data 93%
Similar papers in this journal
- Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis 93%
- Deep in the Bowel: Highly Interpretable Neural Encoder-Decoder Networks Predict Gut Metabolites from Gut Microbiome 93%
- I-Impute: a self-consistent method to impute single cell RNA sequencing data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.