Back

Automatically producing large morphometric datasets from natural history collection images: a case study of Lepidoptera wing shape

Ginot, S.; Debat, V.

2022-07-01 evolutionary biology
10.1101/2022.07.01.497900 bioRxiv
Show abstract

Publicly available image data (2D and 3D) from biological specimens is becoming extremely widespread, notably following digitization efforts from natural history institutions worldwide. To deal with this huge amount of data, high-throughput phenotyping methods are being developed by researchers, to extract biologically meaningful data, in correlation with the burgeoning of the field of phenomics. Here we explore the potential of a combination of simple image treatment algorithms, with a geometric morphometrics contour analysis, applicable to strongly standardized images such as collections of Lepidotera. Using a previously manually landmarked dataset of Morpho butterflies, we show that our automated approach can produce a morphospace similar to that produced by a manual approach. Although the former is more noisy than the latter, it appears to pick up phylogenetic and to some extent ecological signal. Applying then the same approach to a large dataset of images from two different museums, we produce a morphospace containing >5000 specimens, representing 851 species in 24 families of butterflies and moths. The most notable feature of this space is that Sphingidae morphology is clearly separate from the rest, and appears much more constrained. We also show some indirect evidence that at this large interspecific level, potential museum related bias (e.g. inter-user bias in specimen preparation and photography) can be negligible. Altogether, our results suggest that this approach has the potential to produce large-scale analysis of morphology, and could be refined to include more specimens.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.