Back

Evaluating 18S Phylogenetic Placement Accuracy to Uncover Hidden Diversity in Early Branching Animals

Arano-Ansola, J.; Galan-Luque, I.; Demenech, M.; Rico-Martin, L.; Moody, E. R. R.; Suresh, M.; Vaulot, D.; del Campo, J.; Giacomelli, M.; Lozano-Fernandez, J.

2025-12-11 evolutionary biology
10.64898/2025.12.09.693319 bioRxiv
Show abstract

Small subunit ribosomal RNA (SSU rRNA) - 18S in eukaryotes - is a gene universally present across the Tree of Life and it was central in resolving ancient relationships in early molecular phylogenies. Nowadays, despite multi-locus phylogenomic approaches dominate, SSU rRNA sequences still serv as a molecular identifier of biodiversity. Phylogenetic placement uses a reference phylogeny and a given evolutionary model to annotate taxonomically environmental DNA. As such, it enables more accurate identification of divergent sequences that may belong to unknown lineages across the Tree of Life. Here we first created the largest dataset of 18S sequences belonging to non-Bilateria metazoans, by curating the sequences found in PR2 database. We then used it as a case study to test the performance of different 18S-based barcodes (V4, V9, and full-length 18S) within a phylogenetic placement framework. Employing a series of sensitivity analyses, we found that the V9 region generally lacks sufficient phylogenetic signal for reliable placements in most of the cases. The V4 region is accurate when the environmental diversity is well represented in the reference tree but struggles with divergent lineages. Full-length 18S overcome short-read limitations and emerges as the most robust option to uncover unknown major clades. Finally, we apply phylogenetic placement to empirical environmental 18S data. We observe geographical variation in non-bilaterians communities and recover a putative clade of early-branching Ctenophora based on long-sequence barcodes.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.