Integrating multi-OMICS data through sparse Canonical Correlation Analysis for predicting complex traits: A comparative study
Rodosthenous, T.; Evangelou, M.; Shahrezaei, V.
Show abstract
MotivationRecent developments in technology have enabled researchers to collect multiple OMICS datasets for the same individuals. The conventional approach for understanding the relationships between the collected datasets and the complex trait of interest would be through the analysis of each OMIC dataset separately from the rest, or to test for associations between the OMICS datasets. In this work we show that by integrating multiple OMICS datasets together, instead of analysing them separately, improves our understanding of their in-between relationships as well as the predictive accuracy for the tested trait. As OMICS datasets are heterogeneous and high-dimensional (p >> n) integrating them can be done through Sparse Canonical Correlation Analysis (sCCA) that penalises the canonical variables for producing sparse latent variables while achieving maximal correlation between the datasets. Over the last years, a number of approaches for implementing sCCA have been proposed, where they differ on their objective functions, iterative algorithm for obtaining the sparse latent variables and make different assumptions about the original datasets. ResultsThrough a comparative study we have explored the performance of the conventional CCA proposed by Parkhomenko et al. [2009], penalised matrix decomposition CCA proposed by Witten and Tibshirani [2009] and its extension proposed by Suo et al. [2017]. The aferomentioned methods were modified to allow for different penalty functions. Although sCCA is an unsupervised learning approach for understanding of the in-between relationships, we have twisted the problem as a supervised learning one and investigated how the computed latent variables can be used for predicting complex traits. The approaches were extended to allow for multiple (more than two) datasets where the trait was included as one of the input datasets. Both ways have shown improvement over conventional predictive models that include one or multiple datasets. Contacttr1915@ic.ac.uk
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 96%
- Graph Neural Network Modelling as a potentially effective Method for predicting and analyzing Procedures based on Patient Diagnoses 93%
- Fenchel duality of Cox partial likelihood and its application in survival kernel learning 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- RCFGL: Rapid Condition adaptive Fused Graphical Lasso and application to modeling brain region co-expression networks 95%
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 95%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.