Human microbiome sequences in the light of the Nubeam
Dai, H.; Guan, Y.
Show abstract
We present Nubeam (nucleotide be a matrix) as a novel reference-free approach to analyze short sequencing reads. Nubeam represents nucleotides by matrices, transforms a read into a product of matrices, and based on which assigns numbers to reads. Nubeam capitalizes on the non-commutative property of matrix multiplication, such that different reads are assigned different numbers, and similar reads similar numbers. A sample, which is a collection of reads, becomes a collection of numbers that form an empirical distribution. We demonstrate that the genetic difference between samples can be quantified by the distance between empirical distributions. Nubeam can account for GC bias and nucleotide quality, and is computationally efficient; the K-mer method is a special case of Nubeam, but without those benefits. As a reference-free approach, Nubeam avoids reference bias and mapping bias and can work with organisms without reference genomes. Thus, Nubeam is ideal to analyze datasets from metagenomic whole-genome sequencing, where the amount of unmapped reads is substantial. When applied to human microbiome sequencing, Nubeam recapitulated findings made by mapping-based methods, and shed lights on contributions of unmapped reads. In particular, body habitats dominate clustering of unmapped pseudo-samples; there are more outliers in skin whole samples than the skin mapped pseudo-samples; and analysis of unmapped reads suggested that the sequencing depth is far from sufficient for urogenital samples.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- On the impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters 96%
- Dirichlet-multinomial modelling outperforms alternatives for analysis of microbiome and other ecological count data 93%
- On the causes, consequences, and avoidance of PCR duplicates: towards a theory of library complexity 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.