DTH: A nonparametric test for homogeneity of multivariate dispersions
Roy, A.; Satten, G.; Zhao, N.
Show abstract
Testing homogeneity across groups in multivariate data is an important scientific question in its own right, as well as well as an auxiliary step in verifying the assumptions of ANOVA. Existing methods either construct test statistics based on the distance of each observation from the group center, or as the mean of pairwise dissimilarities among observations in a group. Both approaches can fail when mean within-group distance is similar across groups but the distribution of the within-group distances are different. This is a pertinent question in high dimensional microbiome data, where outliers and overdispersion can distort the performance of a mean-dissimilarity-based test. We introduce the non-parametric Distance based Test for Homogeneity (DTH) which measures dissimilarity between groups by comparing the empirical distribution of within-group dissimilarities using a combination of the Kolmogorov-Smirnov and Wasserstein distances. For more than two groups, pairwise group tests are combined using a permutation-based p-value. Through simulations we show that our method has higher power than existing tests for homogeneity in certain situations and comparable power in other situations. We also provide a simple framework for extending the test to a continuous covariate.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Compositional knockoff filter for high-dimensional regression analysis of microbiome data 96%
- A multi-view model for relative and absolute microbial abundances 95%
- A Mixed Effect Similarity Matrix Regression Model (SMRmix) for Integrating Multiple Microbiome Datasets at Community Level and its Application in HIV 93%
Similar papers in this journal
- qad: An R-package to detect asymmetric and directed dependence in bivariate samples 95%
- Robust Bayesian analysis of animal networks subject to biases in sampling intensity and censoring 93%
- A Kernel-Based Change Detection Method to Map Shifts in Phytoplankton Communities Measured by Flow Cytometry 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.