Back

MorphQ: label-free quantification and visualisation of complex morphology from standardised specimen images

Chen, Y.-Y.; Mai, G.-S.; Rubenstein, D. R.; Wei, C.-H.; Shen, S.-F.

2026-08-18 ecology
10.64898/2026.08.11.744091 bioRxiv
Show abstract

O_LIQuantifying complex morphology from images remains difficult because predefined descriptors capture only selected traits. Yet, supervised machine learning models for images require labels and often produce task-specific features that are hard to interpret as biological traits. C_LIO_LIWe present MorphQ, a label-free, self-supervised method that learns a quantitative morphospace from standardised specimen images. Its encoder produces feature vectors for statistical analysis, and its decoder converts analysed positions in morphospace into human-interpretable images, including hypothetical forms not represented by sampled specimens or sampled taxa. C_LIO_LIUsing 1,868 Lepidoptera species, we tested whether MorphQs label-free features were more useful for downstream analysis than features from principal component analysis (PCA) or a supervised species-classification machine learning model. As a diagnostic probe of downstream biological utility, MorphQ features supported higher low-label family-classification accuracy than comparator features, and retained stronger family-level similarity for species absent from model training, indicating better generalisation to species not seen during model training. C_LIO_LITwo case studies link MorphQ morphospaces to species-level elevation and assemblage-level functional diversity while keeping statistical patterns visually inspectable. MorphQ provides a reproducible framework for constructing interpretable morphological trait spaces when predefined descriptors are incomplete and labelled data are limited. C_LI Data/code for peer review: An anonymised repository containing the source code, trained model weights, example data, configuration files and scripts required to reproduce the analyses is available at https://anonymous.4open.science/r/MorphQ-ECD4/.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.