Back

Ridge Redundancy Analysis for High-Dimensional Omics Data

Yoshioka, H.; Aubert, J.; Iwata, H.; Mary-Huard, T.

2025-04-21 bioinformatics
10.1101/2025.04.16.649138 bioRxiv
Show abstract

MotivationRedundancy Analysis (RDA) is a popular reduced rank regression approach for modeling the relationships between two sets of variables. In omics studies, RDA provides a flexible framework for investigating associations between high-dimensional molecular data. However, omics datasets often exhibit strong multicollinearity and suffer from the "large p, small n" problem, where the number of predictor variables p (features) exceeds the number of samples n. Ridge RDA addresses multicollinearity by introducing a ridge penalty in addition to the rank restriction. Despite these advantages, its application to high-dimensional omics data remains challenging due to i) difficulties in manipulating the large coefficient matrix of dimensions p x q, where q represents the number of response variables, and ii) the need to jointly select the rank and the regularization parameter that both influence the performance of the method. ResultsWe propose an efficient computational framework for ridge RDA that overcomes these challenges by leveraging the Singular Value Decomposition of the predictor matrix X. This approach eliminates the need for direct covariance matrix inversion, improving computational efficiency. Furthermore, we introduce a novel strategy that reduces the memory burden associated with the storage of the coefficient matrix from pq to (p + q)r, with r the chosen reduced-rank dimensionality. Our method enables an efficient grid search to jointly select the ridge penalty{lambda} and r via cross-validation. The proposed framework is implemented in the R package rrda, providing a practical and scalable solution for the analysis of high-dimensional omics data. AvailabilityThe R package rrda is available on CRAN: https://CRAN.R-project.org/package=rrda. Scripts and data used for the analysis are available on GitHub: https://github.com/Yoska393/rrda.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.