Correction for spurious taxonomic assignments of k-mer classifiers in low microbial biomass samples using shuffled sequences
Sun, S.; Fodor, A. A.
Show abstract
BackgroundWith the increased use of shotgun metagenome and metatranscriptome sequences in characterizing the microbiome, accurate taxonomic classification of sequencing reads is essential for interpreting microbial community composition and revealing differential microbial signature between groups. K-mer based classifiers such as Kraken2 provide high speed and sensitivity, and are commonly used for low microbial biomass samples. However, their performance can be compromised by specific sources of error without proper parameter settings and incorporation of controls. MethodsIn this study, we analyzed six sequencing datasets of human tumor biopsies with Kraken2 and investigated how shared compact hash codes (i.e., identical hash codes across different k-mers), hash collision and the structure of reference databases can contribute to false positive taxonomic assignments in low biomass samples. ResultsWe demonstrated that in samples with high non-microbial DNA noise, the classified taxa of Kraken2 in sequencing reads are significantly correlated with that of shuffled sequences using the default setting. These taxa showed a similar distribution as those overrepresented in the hash table construction of the reference database. Incorporation of controls using shuffled reads can separate significant taxa with more robust differences from those more affected by background noise. Although the confidence thresholds needed to minimize noise varied with taxa, a minimum value of 0.2 can also help reduce misclassifications. ConclusionOur findings highlighted the need for caution when interpreting low-abundance or unexpected taxa in sequencing datasets of low microbial biomass samples. This work contributes to a more comprehensive understanding of the limitations of k-mer based classification tools and provides practical guidance for improving accuracy in microbiome research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 96%
- Clade-specific long-read sequencing increases the accuracy and specificity of the gyrB phylogenetic marker gene 95%
- Two-target quantitative PCR to predict library composition for shallow shotgun sequencing 95%
Similar papers in this journal
- rRNA Operon Improves Species-Level Classification of Bacteria and Microbial Community Analysis Compared to 16S rRNA 96%
- Library Preparation and Sequencing Platform Introduce Bias in Metagenomic-Based Characterizations of Microbiomes 95%
- Evaluating de novo assembly and binning strategies for time-series drinking water metagenomes. 94%
Similar papers in this journal
- MinION Sequencing of colorectal cancer tumour microbiomes - a comparison with amplicon-based and RNA-Sequencing 95%
- CoSMIC - A hybrid approach for large-scale, high-resolution microbial profiling of novel niches 94%
- Quantitative PCR provides a simple and accessible method for quantitative microbiome profiling 94%
Similar papers in this journal
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 94%
- Meta-analysis of Microbiome Association Networks Reveal Patterns of Dysbiosis in Diseased Microbiomes 94%
- On the robustness of inference of association with the gut microbiota in stool, swab and mucosal tissue samples 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.