Phylogenetic placement and contamination screening of Amoebozoa genomic data from the Protist 10,000 Genomes (P10K) Database
Porfirio-Sousa, A. L.; Jones, R. E.; Brown, M. W.; Lahr, D. J. G.; Tice, A. K.
Show abstract
BackgroundGenomic data are essential for uncovering the evolutionary history, ecological roles, and diversity of life. Yet, microbial eukaryotes like Amoebozoa, an ancient and morphologically diverse lineage, remain critically underrepresented in genomic repositories. This has limited our ability to address fundamental questions in eukaryotic evolution. The Protist 10,000 Genomes (P10K) initiative seeks to fill this gap by generating and compiling genome- and transcriptome-level data for a wide range of microbial eukaryotes. To ensure the reliability of these resources, accurate taxonomic identification and contamination screening are vital. In this study, we aimed to assess the taxonomic consistency and integrity of the P10K database with a phylogenetic-based approach using Amoebozoa as a case study. ResultsThrough SSU rDNA/rRNA and COI phylogenetic reconstructions this study confirmed several initial taxonomic identifications provided in the P10K database, resolved ambiguities at higher taxonomic levels, and corrected misassignments among morphologically similar but phylogenetically distant taxa. Moreover, the contamination screening using SSU rDNA/rRNA revealed several amoebozoan data that are contaminated by sequence from other eukaryotic taxa, representing contaminated genomic assemblies. ConclusionPhylogenetic placement coupled with contamination screening enabled us to distinguish the higher-quality Amoebozoa datasets currently available in the P10K database from those requiring decontamination or additional sequencing before downstream use. These findings serve as a reference for the future use of these data and as a guide for further sequencing efforts aimed at expanding the taxonomic diversity of Amoebozoa represented at the genomic level. By applying a phylogenetic survey to the Amoebozoa data, we present a framework that can be extended to other microbial eukaryote lineages. Addressing imprecise taxonomic identifications and contamination in certain P10K datasets, as well as data reproducibility, will further enhance the value of this unprecedented genomic resource for protists, with significant potential to illuminate the evolution and diversification of eukaryotic life.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond the limits of the unassigned protist microbiome: inferring large-scale spatio-temporal patterns of marine parasites 95%
- Discovery and genomics of H2-oxidizing/O2-reducing Deferribacterota ectosymbiotic with protists in the guts of termites and a Cryptocercus cockroach 95%
- Ecology and evolution of chlamydial symbionts of arthropods 94%
Similar papers in this journal
- FANGORN: A quality-checked and publicly available database of full-length 16S-ITS-23S rRNA operon sequences 95%
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 94%
- Delineating bacterial genera based on gene content analysis: a case study of the Mycoplasmatales-Entomoplasmatales clade within the class Mollicutes 94%
Similar papers in this journal
- Utilisation of Oxford Nanopore sequencing to generate six complete gastropod mitochondrial genomes as part of a biodiversity curriculum 95%
- Extreme mito-nuclear discordance within Anthozoa, with notes on unique properties of their mitochondrial genomes 94%
- Genomic analysis of Coccomyxa viridis, a common low-abundance alga associated with lichen symbioses 94%
Similar papers in this journal
- A Small Genome Amidst the Giants: Evidence of Genome Reduction in a Small Tubulinid Free-Living Amoeba 93%
- A chromosome-level assembly and functional genomic resources for the model annelid Capitella teleta 93%
- The genome assembly of Rhabditoides inermis from a complex microbial community reveals further evidence for parallel gene family expansions across multiple nematodes 93%
Similar papers in this journal
- A flexible pipeline combining clustering and correction tools for prokaryotic and eukaryotic metabarcoding 96%
- Validated removal of nuclear pseudogenes and sequencing artefacts from mitochondrial metabarcode data 94%
- The Ribosomal Operon Database (ROD): A full-length rDNA operon database extracted from genome assemblies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.