Back

Discovering the unseen: a performance comparison of taxonomic classification methods under unknown DNA barcodes

Orsholm, J.; Zito, A.; Somervuo, P.; Harrison, J. P.; Koskela, M.; Ovaskainen, O.; Braga, M. P.; Chazot, N.; Roslin, T.; Furneaux, B.

2025-11-24 bioinformatics
10.1101/2025.10.13.681976 bioRxiv
Show abstract

O_LIDNA barcoding and metabarcoding have emerged as cost-efficient, standardized methods for characterizing local biodiversity. Based on the sequencing of a small targeted gene fragment, it is theoretically possible to identify a wide diversity of taxa by comparing them with reference sequence databases. However, a key challenge for accurate taxonomic classification is the incompleteness of such databases, leading to most query sequences lacking species-level matches. C_LIO_LIWhere species-level matches are missing, it may be possible to classify query sequences to a higher taxonomic level, such as genus or family, based on the similarity of related reference taxa. The challenge then lies in confidently recognizing whether the sequence belongs to an unobserved (here, "novel") taxon on a given taxonomic level. C_LIO_LIIn this study, we evaluated the performance and utility of several methods for taxonomic classification. Methods were assessed based on the classification accuracy of both observed and novel taxa, training time, space requirements, and run time. We did this for two cases: the COI barcode for arthropods, and the ITS barcode for fungi, with the latter representing an instance with substantially greater sequence similarity variation within classes. To test classification of novel taxa, we used well-curated datasets with partially distinct taxonomic distribution between the training and test set. Novel taxa occurred at all evaluated taxonomic levels, such as novel species in observed genera and novel genera in observed families. We further assessed the effect on performance when shifting from full-length barcodes to shorter sequences as generated through metabarcoding in the testing dataset. C_LIO_LIThis study sheds light on the strengths and limitations of different classification algorithms across varied ecological contexts and provides guidance for researchers in selecting suitable algorithms for DNA barcoding and metabarcoding applications. In particular, it demonstrates the supreme performance of phylogenetic placement methods such as EPA-ng for classification of arthropod COI barcodes, and composition-based classifiers such as SINTAX, RDP-NBC, and IDTAXA for fungal ITS. C_LI

Published in Methods in Ecology and Evolution (predicted rank #4) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.