Fast and interpretable quantification of biological shape heterogeneity via stratified Wasserstein kernel
Zhao, W.; Sutherland, D. J.; Dao Duc, K.
Show abstract
Modern imaging technologies produce vast collections of cellular and subcellular structures, calling for principled methods that enable shape comparison across individuals and populations. We introduce the stratified Wasserstein framework, which treats each shape as an unstructured point cloud and embeds it into Euclidean space via ranked local distance profiles. This embedding yields an isometry-invariant Euclidean distance and a positive-definite kernel for population analysis, with a consistent sample-based estimator that supports large datasets in near-quadratic time. By leveraging kernel methods, the framework enables statistically rigorous tasks such as nonparametric hypothesis testing, providing theoretical guarantees as well as interpretability. We demonstrate our frameworks applicability to large-scale biological datasets. Analyzing 2D cancer cell contours, we quantify population-level discrepancies and identify representative cells contributing most strongly to the observed differences. Using 3D volumes of cell envelope and nucleus, we reveal progression patterns that capture morphological changes across cell populations both at the level of individual shapes. These results establish a simple and principled tool for population-level biological shape analysis, with potential impact across diverse domains of computational imaging and data science.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 96%
- A comprehensive comparison on cell type composition inference for spatial transcriptomics data 93%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.