Clustering using Automatic Differentiation: Misspecified models and High-Dimensional Data
Kasa, S. R.; Rajan, V.
Show abstract
We study two practically important cases of model based clustering using Gaussian Mixture Models: (1) when there is misspecification and (2) on high dimensional data, in the light of recent advances in Gradient Descent (GD) based optimization using Automatic Differentiation (AD). Our simulation studies show that EM has better clustering performance, measured by Adjusted Rand Index, compared to GD in cases of misspecification, whereas on high dimensional data GD outperforms EM. We observe that both with EM and GD there are many solutions with high likelihood but poor cluster interpretation. To address this problem we design a new penalty term for the likelihood based on the Kullback Leibler divergence between pairs of fitted components. Closed form expressions for the gradients of this penalized likelihood are difficult to derive but AD can be done effortlessly, illustrating the advantage of AD-based optimization. Extensions of this penalty for high dimensional data and for model selection are discussed. Numerical experiments on synthetic and real datasets demonstrate the efficacy of clustering using the proposed penalized likelihood approach.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EpiLPS: a fast and flexible Bayesian tool for near real-time estimation of the time-varying reproduction number 95%
- Efficient inference for agent-based models of real-world phenomena 94%
- CrossLabFit: A Novel Framework for Integrating Qualitative and Quantitative Data Across Multiple Labs for Model Calibration 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.