Stability and Performance of Linear Combination Tests of Gene Set Enrichment for Multiple Covariance Estimators in Unbalanced Studies
Khademioureh, S.; Amini, P.; Ghasemi, E.; Calistrate-Petre, P.; Pyne, S.; Dinu, I.
Show abstract
Gene set analysis (GSA) is essential for understanding coordinated gene expression changes within biological pathways, especially in high-dimensional data generated by platforms such as RNA-seq and microarrays. This study focuses on the linear combination test (LCT), a GSA method that combines multiple genelevel statistics into a powerful test statistic to assess the association between a gene set and outcomes of interest in a given set of samples. We evaluated the performance and stability of LCT using different covariance matrix estimators, including ridge, graphical lasso, and adaptive lasso, which are known for their effectiveness in high-dimensional data analysis. In addition, we assessed the robustness of LCT in the face of unbalanced study designs, which are typical in biomedical research due to limited sample availability and the high cost of data generation. We conducted a simulation study and applied LCT to publicly available gene expression datasets comparing patients with systemic lupus erythematosus (SLE) to healthy controls, where the number of controls is significantly lower than the number of cases. Our findings demonstrate that while LCTs default shrinkage estimator shows limitations in highly correlated and unbalanced designs, ridge estimation provides a more reliable alternative for unbalanced scenarios. Researchers can optimize LCTs performance by selecting appropriate covariance estimators based on their data structure. These results suggest that LCT is a reliable and powerful tool for GSA in unbalanced studies, identifying SLE-relevant gene sets more effectively than other GSA methods and showing validation against clinical phenotypes, offering valuable insights into the underlying mechanisms of complex diseases such as SLE.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A robust and fast two-sample test of equal correlations with an application to differential co-expression 96%
- Using a supervised principal components analysis for variable selection in high-dimensional datasets reduces false discovery rates 95%
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.