Back

Extending and improving metagenomic taxonomic profiling with uncharacterized species with MetaPhlAn 4

Blanco-Miguez, A.; Beghini, F.; Cumbo, F.; McIver, L. J.; Thompson, K. N.; Zolfo, M.; Manghi, P.; Dubois, L.; Huang, K. D.; Thomas, A. M.; Piccinno, G.; Piperni, E.; Puncochar, M.; Valles-Colomer, M.; Tett, A.; Giordano, F.; Davies, R.; Wolf, J.; Berry, S. E.; Spector, T. D.; Franzosa, E. A.; Pasolli, E.; Asnicar, F.; Huttenhower, C.; Segata, N.

2022-08-22 bioinformatics
10.1101/2022.08.22.504593 bioRxiv
Show abstract

Metagenomic assembly enables novel organism discovery from microbial communities, but from most metagenomes it can only capture few abundant organisms. Here, we present a method - MetaPhlAn 4 - to integrate information from both metagenome assemblies and microbial isolate genomes for improved and more comprehensive metagenomic taxonomic profiling. From a curated collection of 1.01M prokaryotic reference and metagenome-assembled genomes, we defined unique marker genes for 26,970 species-level genome bins, 4,992 of them taxonomically unidentified at the species level. MetaPhlAn 4 explains [~]20% more reads in most international human gut microbiomes and >40% in less-characterized environments such as the rumen microbiome, and proved more accurate than available alternatives on synthetic evaluations while also reliably quantifying organisms with no cultured isolates. Application of the method to >24,500 metagenomes highlighted previously undetected species to be strong biomarkers for host conditions and lifestyles in human and mice microbiomes, and showed that even previously uncharacterized species can be genetically profiled at the resolution of single microbial strains. MetaPhlAn 4 thus integrates the novelty of metagenomic assemblies with the sensitivity and fidelity of reference-based analyses, providing efficient metagenomic profiling of uncharacterized species and enabling deeper and more comprehensive microbiome biomarker detection.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.