Back

Multiset partial least squares with rank order of groups for integrating multi-omics data

Yamamoto, H.

2022-08-31 bioinformatics
10.1101/2022.08.30.505949 bioRxiv
Show abstract

As more multi-omics data such as metabolome, proteome and transcriptome data are acquired, statistical analyses to integrate multi-omics data has been required. Multivariate analysis for integration of multi-omics data has been proposed such as multiset partial least squares (PLS). However, when we want to extract scores about the rank order of groups such as severity of disease as well as difference of groups, we could not always extract scores with rank order of groups by using multiset PLS. Therefore, we proposed multiset PLS with rank order of groups (multiset PLS-ROG) to integrate multiple omics dataset. We visualized multiple omics data by using multiset PLS-ROG in two studies of multi-omics and multi-organ derived metabolome data. After we visualize data and focus on a specific score with phenotype of interest, it is important to select compounds to identify biomarker candidates or make biological inferences. In ordinary principal component analysis (PCA) or PLS, we could select statistical significantly compounds correlated with PC or PLS score by using PC or PLS loadings. However, multiset PLS loading has not been used because its statistical property has not been clarified. Then, we clarified statistical property of multiset PLS and multiset PLS-ROG loading, and we could identify statistically significant compounds by using statistical hypothesis testing. We implemented multiset PLS-ROG to R loadings package, freely available in CRAN.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.