Discovering the unseen: a performance comparison of taxonomic classification methods under unknown DNA barcodes
Orsholm, J.; Zito, A.; Somervuo, P.; Harrison, J. P.; Koskela, M.; Ovaskainen, O.; Braga, M. P.; Chazot, N.; Roslin, T.; Furneaux, B.
Show abstract
O_LIDNA barcoding and metabarcoding have emerged as cost-efficient, standardized methods for characterizing local biodiversity. Based on the sequencing of a small targeted gene fragment, it is theoretically possible to identify a wide diversity of taxa by comparing them with reference sequence databases. However, a key challenge for accurate taxonomic classification is the incompleteness of such databases, leading to most query sequences lacking species-level matches. C_LIO_LIWhere species-level matches are missing, it may be possible to classify query sequences to a higher taxonomic level, such as genus or family, based on the similarity of related reference taxa. The challenge then lies in confidently recognizing whether the sequence belongs to an unobserved (here, "novel") taxon on a given taxonomic level. C_LIO_LIIn this study, we evaluated the performance and utility of several methods for taxonomic classification. Methods were assessed based on the classification accuracy of both observed and novel taxa, training time, space requirements, and run time. We did this for two cases: the COI barcode for arthropods, and the ITS barcode for fungi, with the latter representing an instance with substantially greater sequence similarity variation within classes. To test classification of novel taxa, we used well-curated datasets with partially distinct taxonomic distribution between the training and test set. Novel taxa occurred at all evaluated taxonomic levels, such as novel species in observed genera and novel genera in observed families. We further assessed the effect on performance when shifting from full-length barcodes to shorter sequences as generated through metabarcoding in the testing dataset. C_LIO_LIThis study sheds light on the strengths and limitations of different classification algorithms across varied ecological contexts and provides guidance for researchers in selecting suitable algorithms for DNA barcoding and metabarcoding applications. In particular, it demonstrates the supreme performance of phylogenetic placement methods such as EPA-ng for classification of arthropod COI barcodes, and composition-based classifiers such as SINTAX, RDP-NBC, and IDTAXA for fungal ITS. C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Tackling the phylogenetic conundrum of Hydroidolina (Cnidaria: Medusozoa: Hydrozoa) by assessing competing tree topologies with targeted high-throughput sequencing. 94%
- Robust genome-based delineation of bacterial genera 94%
- Mitochondrial genomes of Columbicola feather lice are highly fragmented, indicating repeated evolution of minicircle-type genomes in parasitic lice 93%
Similar papers in this journal
Similar papers in this journal
- Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories 95%
- Whole-genomes illuminate the drivers of gene tree discordance and the tempo of tinamou diversification (Aves: Tinamidae) 95%
- Genomic characterization and curation of UCEs improves species tree reconstruction. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.