Benchmarking Phenotypic Clustering Algorithms via Empirically Calibrated Simulations: A Diagnostic Framework to Improve Biodiversity Assessment in Neglected Crops
NAINO JIKA, A. K.
Show abstract
Clustering algorithms are widely used for phenotypic characterization and germplasm management, particularly in neglected and underutilized species (NUS) that lack genomic resources. However, their performance under biologically realistic conditions remains poorly understood. Standard clustering methods commonly applied in crop research often assume distinct, isotropic, and homogeneous clusters--assumptions rarely satisfied in real-world NUS datasets. We developed a biologically informed simulation framework, empirically calibrated with phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios. Our simulations integrated heterogeneous trait distributions (normal, gamma), strong inter-trait correlations (up to r = -0.84), heteroscedasticity, and moderate population structure (Pst {approx} 0.15), as observed in fonio landraces. Each scenario was replicated 100 times, with clustering accuracy evaluated using Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), Silhouette coefficient, and Davies-Bouldin index. The results revealed consistently poor algorithm performance under realistic conditions (e.g., ARI < 0.07), including for widely used methods in NUS research such as K-means, GMM, and PAM. Performance markedly improved under idealized conditions, validating our simulation framework. These findings highlight the risk of overinterpreting clustering outputs from weakly structured phenotypic datasets and expose key limitations in current biodiversity analysis practices--particularly those guiding plant genetic resource conservation programs. We provide an open-source R-based diagnostic tool, available on Zenodo 5(https://doi.org/10.5281/zenodo.15877863), to assist practitioners in selecting robust clustering approaches for germplasm management and pre-breeding in data-scarce crops.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Automated SNP-Based Approach for Contaminant Identification in Biparental Polyploid Populations of Tropical Forage Grasses 94%
- Enviromic assembly increases accuracy and reduces costs of the genomic prediction for yield plasticity 93%
- Incorporating gene expression and environment improves genomic prediction of wheat traits 92%
Similar papers in this journal
- Back to the future 2: the implications of germplasm structure on thebalance between short and long-term genetic gain in a changingtarget population of environments 94%
- PyBrOpS: a Python package for breeding program simulation and optimization for multi-objective breeding 93%
- Improved Genomic Prediction Performance with Ensembles of Diverse Models 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.