Back

Benchmarking Phenotypic Clustering Algorithms via Empirically Calibrated Simulations: A Diagnostic Framework to Improve Biodiversity Assessment in Neglected Crops

NAINO JIKA, A. K.

2025-07-21 plant biology
10.1101/2025.07.16.665063 bioRxiv
Show abstract

Clustering algorithms are widely used for phenotypic characterization and germplasm management, particularly in neglected and underutilized species (NUS) that lack genomic resources. However, their performance under biologically realistic conditions remains poorly understood. Standard clustering methods commonly applied in crop research often assume distinct, isotropic, and homogeneous clusters--assumptions rarely satisfied in real-world NUS datasets. We developed a biologically informed simulation framework, empirically calibrated with phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios. Our simulations integrated heterogeneous trait distributions (normal, gamma), strong inter-trait correlations (up to r = -0.84), heteroscedasticity, and moderate population structure (Pst {approx} 0.15), as observed in fonio landraces. Each scenario was replicated 100 times, with clustering accuracy evaluated using Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), Silhouette coefficient, and Davies-Bouldin index. The results revealed consistently poor algorithm performance under realistic conditions (e.g., ARI < 0.07), including for widely used methods in NUS research such as K-means, GMM, and PAM. Performance markedly improved under idealized conditions, validating our simulation framework. These findings highlight the risk of overinterpreting clustering outputs from weakly structured phenotypic datasets and expose key limitations in current biodiversity analysis practices--particularly those guiding plant genetic resource conservation programs. We provide an open-source R-based diagnostic tool, available on Zenodo 5(https://doi.org/10.5281/zenodo.15877863), to assist practitioners in selecting robust clustering approaches for germplasm management and pre-breeding in data-scarce crops.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.