Back

Biome-specific genome catalogues reveal functional potential of shallow sequencing

Escobar-Zepeda, A.; Ruuskanen, M. O.; Beracochea, M.; Lu, J.; Mongad, D.; Richardson, L.; Finn, R. D.; Lahti, L.

2025-06-20 bioinformatics
10.1101/2025.06.16.659887 bioRxiv
Show abstract

The use of 16S rRNA metabarcoding for functional prediction is limited by several biases. Shallow shotgun sequencing is a cost-effective and taxonomically high-resolution alternative to 16S rRNA metabarcoding, but the low sequencing depth limits functional inference. Our BioSIFTR tool maps shallow shotgun sequencing reads against single-biome databases and extrapolates into their precalculated functional profiles. We used three datasets from red junglefowl, mice, and human gut, containing matched deep shotgun metagenomic and 16S rRNA metabarcoding data for taxonomic and functional benchmarking. An additional human gut deep shotgun sequencing data set was subsampled to 1 M reads and analysed with BioSIFTR to replicate previously obtained results. BioSIFTR taxonomic and functional profiles closely agree with the results of the full deep sequencing data in all biomes. We also replicated differences in the human gut microbiome between high and low trimethylamine N-oxide producing participants, using only < 2 % of the original deep sequencing data. The BioSIFTR tool is a powerful approach which approximates the functional information of a deep-sequenced metagenome while using only a fraction of the data. Shallow shotgun sequencing combined with BioSIFTR could be a stand-in replacement for 16S rRNA metabarcoding with an increased taxonomic and functional resolution, and lower bias. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=87 SRC="FIGDIR/small/659887v1_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@8ca365org.highwire.dtl.DTLVardef@13b466dorg.highwire.dtl.DTLVardef@8cad19org.highwire.dtl.DTLVardef@1b853fa_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.