Removal of rare amplicon sequence variants from 16S rRNA gene sequence surveys biases the interpretation of community structure data
Schloss, P. D.
Show abstract
Methods for remediating PCR and sequencing artifacts in 16S rRNA gene sequence collections are in continuous development and have significant ramifications on the inferences that can be drawn. A common approach is to remove rare amplcon sequence variants (ASVs) from datasets. But, the definition of rarity is generally selected without regard for the number of sequences in the samples or the variation in sequencing depth across samples within a study. I analyzed the impact of removing rare ASVs on metrics of alpha and beta diversity using samples collected across 12 published datasets. Removal of rare ASVs significantly decreased the number of ASVs and operational taxonomic units as well as their diversity. Furthermore, their removal increased the variation in community structure between samples. When simulating a known effect size, removal of rare ASVs reduced the power to detect the effect relative to not removing rare ASVs. Removal of rare ASVs did not affect the false detection rate when samples were randomized to simulate a null model. However, the false detection rate increased when rare ASVs were removed using a null distribution and assignment of samples to simulated treatment groups according to their sequencing depth. The false detection rate did not vary when rare ASVs were retained. This analysis demonstrates the problems inherent in removing rare ASVs. Researchers are encouraged to retain rare ASVs, to select approaches that minimize PCR and sequencing artifacts, and to use rarefaction to control for uneven sequencing effort. ImportanceRemoving rare amplicon sequence variants (ASVs) from 16S rRNA gene sequence collections is an approach that has grown in popularity for limiting PCR and sequencing artifacts. Yet, it is unclear what impact an abundance-based filter has on downstream analyses. To investigate the effects of removing rare ASVs, I analyzed the community distributions found in the samples of 12 published datasets. Analysis of these data and simulations based on them showed that removal of rare ASVs distorts the representation of microbial communities. This has the effect of artificially making it more difficult to detect differences between treatment groups. Also of concern was the observation that if sequencing depth is confounded with the treatment, then the probability of falsely detecting a difference between the treatment groups increased with the removal of rare ASVs. The practice of removing rare ASVs should stop, lest researcher adversely affect the interpretation of their data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Waste not, want not: Revisiting the analysis that called into question the practice of rarefaction 96%
- Amplicon sequence variants artificially split bacterial genomes into separate clusters 95%
- Ecological Observations Based on Functional Gene Sequencing Are Sensitive to the Amplicon Processing Method 94%
Similar papers in this journal
Similar papers in this journal
- Evaluation of the effects of library preparation procedure and sample characteristics on the accuracy of metagenomic profiles 96%
- A standardized quantitative analysis strategy for stable isotope probing metagenomics 95%
- Adapting macroecology to microbiology: using occupancy modelling to assess functional profiles across metagenomes 94%
Similar papers in this journal
- Natural bacterial assemblages in Arabidopsis thaliana tissues become more distinguishable and diverse during host development 95%
- Inferring species compositions of complex fungal communities from long- and short-read sequence data 93%
- Microbial populations are shaped by dispersal and recombination in a low biomass subseafloor habitat 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.