Back

Integrating multi-OMICS data through sparse Canonical Correlation Analysis for predicting complex traits: A comparative study

Rodosthenous, T.; Evangelou, M.; Shahrezaei, V.

2019-11-15 genomics
10.1101/843524 bioRxiv
Show abstract

MotivationRecent developments in technology have enabled researchers to collect multiple OMICS datasets for the same individuals. The conventional approach for understanding the relationships between the collected datasets and the complex trait of interest would be through the analysis of each OMIC dataset separately from the rest, or to test for associations between the OMICS datasets. In this work we show that by integrating multiple OMICS datasets together, instead of analysing them separately, improves our understanding of their in-between relationships as well as the predictive accuracy for the tested trait. As OMICS datasets are heterogeneous and high-dimensional (p >> n) integrating them can be done through Sparse Canonical Correlation Analysis (sCCA) that penalises the canonical variables for producing sparse latent variables while achieving maximal correlation between the datasets. Over the last years, a number of approaches for implementing sCCA have been proposed, where they differ on their objective functions, iterative algorithm for obtaining the sparse latent variables and make different assumptions about the original datasets. ResultsThrough a comparative study we have explored the performance of the conventional CCA proposed by Parkhomenko et al. [2009], penalised matrix decomposition CCA proposed by Witten and Tibshirani [2009] and its extension proposed by Suo et al. [2017]. The aferomentioned methods were modified to allow for different penalty functions. Although sCCA is an unsupervised learning approach for understanding of the in-between relationships, we have twisted the problem as a supervised learning one and investigated how the computed latent variables can be used for predicting complex traits. The approaches were extended to allow for multiple (more than two) datasets where the trait was included as one of the input datasets. Both ways have shown improvement over conventional predictive models that include one or multiple datasets. Contacttr1915@ic.ac.uk

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
15.3%
2
BMC Bioinformatics
457 papers in training set
Top 0.6%
11.2%
3
Bioinformatics
1204 papers in training set
Top 3%
8.0%
4
Briefings in Bioinformatics
354 papers in training set
Top 0.8%
8.0%
5
PLOS Computational Biology
1863 papers in training set
Top 5%
6.8%
6
PLOS ONE
5266 papers in training set
Top 31%
4.9%
50% of probability mass above
7
Frontiers in Genetics
230 papers in training set
Top 0.8%
4.1%
8
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.7%
9
Scientific Reports
3612 papers in training set
Top 42%
2.5%
10
The Annals of Applied Statistics
19 papers in training set
Top 0.1%
2.4%
11
BMC Genomics
406 papers in training set
Top 4%
2.1%
12
Biostatistics
24 papers in training set
Top 0.2%
1.7%
13
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
14
GigaScience
212 papers in training set
Top 2%
1.7%
15
Genes
144 papers in training set
Top 2%
1.5%
16
Journal of Computational Biology
48 papers in training set
Top 0.8%
1.1%
17
Biometrics
23 papers in training set
Top 0.2%
1.1%
18
Statistical Methods in Medical Research
11 papers in training set
Top 0.1%
1.0%
19
Computers in Biology and Medicine
128 papers in training set
Top 4%
1.0%
20
Statistics in Medicine
40 papers in training set
Top 0.5%
0.9%
21
Journal of Mathematical Biology
40 papers in training set
Top 0.5%
0.9%
22
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.9%
23
Mathematical Biosciences
49 papers in training set
Top 1%
0.6%
24
PeerJ
308 papers in training set
Top 13%
0.6%
25
Heliyon
152 papers in training set
Top 9%
0.6%