Permutation analysis prior to variable selection greatly enhances robustness of OPLS analysis in small cohorts
Ström, M.; Wheelock, A. M.
Show abstract
The R-workflow ropls-ViPerSNet (R orthogonal projections of latent structures with Variable Permutation Selection and Elastic Net) facilitates variable selection, model optimization and significance testing using permutations of OPLS-DA models, with the scaled loadings (p[corr]) as the main metric of significance cutoff. Permutations including (over) the variable selection procedure, prior to (pre-), as well as post variable selection are performed. The resulting p-values for the correlation of the model (R2) and the cross-validated correlation of the model (Q2) pre-, post- and over-variable selection are provided as additional model statistics. These model statistics are useful for determining the true significance level of OPLS models, which otherwise have proven difficult to assess particularly for small sample sizes. Furthermore, a means for estimating the background noise level based on permuted false positive rates of R2 and Q2 is proposed. This novel metric is then used to calculate an adjusted Q2 value. Using a publicly available metabolomics dataset, the advantage of performing permutations over variable selection was demonstrated for small sample sizes. Iteratively reducing the sample sizes resulted in overinflated models with increasing R2 and Q2 and permutations post variable selection indicated falsely significant models. In contrast, the adjusted Q2 was marginally affected by sample size, and represents a robust estimate of model predictability, and permutations over variable selection showed true significance of the models. An additional Elastic Net variable selection option is included in the workflow for variable selection by coefficient value penalization using an iterative approach to reduce noise while avoiding overfitting.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Repeated measures ASCA+ for analysis of longitudinal intervention studies with multivariate outcome data 94%
- PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes 94%
- Sparse Multitask group Lasso for Genome-Wide Association Studies 92%
Similar papers in this journal
- ACMTF-R: supervised multi-omics data integration uncovering shared and distinct outcome-associated variation 94%
- Interpretable machine learning with tree-based Shapley additive explanations: application to metabolomics datasets for binary classification 93%
- OpenStats: A Robust and Scalable Software Package for Reproducible Analysis of High-Throughput Phenotypic Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.