An ensemble learning method for joint kernel association testing and principal component analysis on multiple kernels
Koh, H.
Show abstract
In high-dimensional omics studies, researchers often conduct kernel association testing to power-fully detect the relationship of the genetic or microbial composition with human health or disease. Especially, in human microbiome studies, its dimension reduction analysis follows to visually represent complex microbiome data in a simple two- or three-dimensional coordinate space. However, various kernels exist, and they produce all different outcomes; hence, it is hard to interpret them all consistently. Then, omnibus testing has recently been a subject of intense investigation for a unified and powerful statistical inference. However, current omnibus tests are purely a test for significance producing only a P-value as their outcome with no related dimension reduction and visualization approach; hence, their utility is still limited. In this paper, I introduce an ensemble learning method, named as enKern, for joint kernel associating testing and principal component analysis on multiple kernels. enKern is based on a weight learning scheme that leverages complementary contributions from multiple kernels for powerful performance for various association patterns. I show that applying the weights to individual test statistics or individual kernels is equivalent, which in turn enables a visualization in a reduced dimensional coordinate space based on the weighted kernel to be matched with its original significance testing scheme. I demonstrate its use for human microbiome {beta}-diversity analysis. I also demonstrate its outperformance in validity and power through simulation experiments. enKern is freely available in R computing environment at https://github.com/hk1785/enkern.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Dirichlet distribution parameter estimation with application in microbiome analyses 96%
- A robust and fast two-sample test of equal correlations with an application to differential co-expression 94%
- Using a supervised principal components analysis for variable selection in high-dimensional datasets reduces false discovery rates 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.