Back

Comprehensive Metagenomic Profiling of Diverse Microbiomes

Kakuk, B.; Dormo, A.; Taifi, A.; Jaray, T.; Kurucsai, G.; Gulyas, G.; Prazsak, I.; Boldogkoi, Z.; Tombacz, D.

2025-02-13 microbiology
10.1101/2025.02.13.638056 bioRxiv
Show abstract

The canine gut microbiome serves as a key model for veterinary and human health research, but inconsistent findings arise due to methodological variations. This study presents a three-part dataset to clarify how DNA extraction, primer selection, and sequencing platforms influence microbial profiling. First, we performed ultra-deep sequencing of a single dog fecal sample using five DNA isolation kits, multiple library protocols, and four sequencing platforms (Illumina MiSeq/NovaSeq, ONT MinION, PacBio Sequel IIe), enabling direct comparisons of 16S rRNA and shotgun sequencing techniques. Second, we analyzed 40 fecal samples from eight co-housed dogs using Zymo High-Molecular-Weight (ZHMW) and Zymo MagBead (ZMB) extraction kits to assess longitudinal extraction-kit effects. Third, we evaluated three 16S primer systems (standard ONT, PacBio, and modified ONT with degenerate bases) using synthetic mock communities and human/canine fecal samples to quantify primer biases. By integrating synthetic and biological replicates, this dataset provides a standardized resource for benchmarking bioinformatics pipelines and improving cross-study comparability. The study generated 75.31GB of new sequencing data: 43.451GB from ZHMW-ZMB comparisons, 22.611GB for primer evaluations, and 9.191GB from the single-sample analysis. Combined with 31.51GB of prior data, the total dataset exceeds 1061GB, including all analytical outputs. These resources enhance methodological transparency and accuracy in canine gut microbiome research across diverse laboratory workflows.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.