Contamination of widely used databases compromises rRNA based taxonomic assignment of metagenomes and metatranscriptomes
Grant, A.; Davies, C. S.
Show abstract
Taxonomic annotation of metagenomic and metatranscriptomic datasets frequently relies on aligning small subunit (SSU) rRNA reads to annotated reference databases. However, widely used SSU databases, including SILVA, GTDB, Eukaryome, and those distributed with SortMeRNA, are contaminated with large subunit (LSU) rRNA sequences. Low levels of LSU contamination can lead to large numbers of LSU reads being incorrectly identified as SSU and have serious consequences for downstream taxonomic assignments. The problem can be avoided by rigorously removing LSU sequences from databases used for taxonomic annotation, which we have done for KSGP 4.0. Alternatively, initial SSU read selection can be carried out with a carefully curated small SSU database that is free of LSU contamination. We illustrate these two approaches in combination using a metatranscriptomic dataset from an estuarine sediment. Without database cleaning, LSU-derived reads can make up half of supposed SSU sequences and are assigned to a small number of apparently dominant but artefactual taxa. Database cleaning removes this problem and we provide a script, total_rnaseq, to annotate total RNASeq or metagenomic data using this approach. The KSGP 4.0 database provides substantially improved annotation of Archaea compared to the SILVA database and moderate and small improvements for eukaryotes and bacteria respectively.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 93%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 93%
- Diversity and distribution of sediment bacteria across an ecological and trophic gradient 92%
Similar papers in this journal
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 95%
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 94%
- FANGORN: A quality-checked and publicly available database of full-length 16S-ITS-23S rRNA operon sequences 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Ecological Observations Based on Functional Gene Sequencing Are Sensitive to the Amplicon Processing Method 95%
- Waste not, want not: Revisiting the analysis that called into question the practice of rarefaction 94%
- Amplicon sequence variants artificially split bacterial genomes into separate clusters 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.