Back

Integrity and miss grouping as support for clusters in agglomerative hierarchical methods: the R-package octopucs

MacGregor-Fors, I.; Guevara, R.

2024-08-05 bioinformatics
10.1101/2024.08.01.606070 bioRxiv
Show abstract

The hierarchical clustering of communities based on speciescompositional similarity (and abundance or frequency) is standard in community ecology to unveil large-scale patterns and underlay environmental causes of differentiation among communities. Often, the threshold to discretize clusters is arbitrary despite the existence of methods that minimize this bias. Most available techniques use the exact repeatability of clusters memberships under resampling protocols to define robust groups. Here, we propose a novel method to yield cluster support throughout the topology of hierarchical analyses. We acknowledge that the observed dataset may be biased. Instead of using the observed topology as a reference to work out the groups support, we compiled a consensus topology. Then, we borrowed the ecological concepts of reciprocal complementarities between a pair of communities and translated them into cluster integrity and contamination. This procedure allows for building support for groups even when there is a partial membership match after resampling the dataset. In addition, we present the R package octopucs in which we implemented the method reported here. Compared with other methods, the new proposal robustly detected changes in the group memberships, resulting in considerable differences in the pattern of supported clusters.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.