MGM2 as a Unified Foundation Model for Microbiome World Exploration
Zhang, H.; Zhang, Y.; Qi, Y.; Liu, T.; Yang, R.; Ning, K.
Show abstract
Microbiomes are information-rich biological systems, yet most computational analyses still reduce communities to cohort-specific abundance tables. Here we introduce MGM2, a multimodal foundation model pretrained on 1,821,291 MicrobeAtlas samples and 225,067 OTUs clustered at 99% sequence similarity. MGM2 couples NTv3-derived microbial sequence embeddings with abundance conditioning and community-semantic alignment to learn transferable sample- and token-level representations. Frozen MGM2 representations outperformed DeepPhylo by 0.06-0.21 macro-AUROC across five temporally held-out MGnify hierarchy levels, with the largest gains for rare and fine-grained labels. In fecal microbiota transplantation, MGM2-XLarge achieved a response ROC AUC of 0.79 and reduced post-transplant Bray-Curtis distance by 15% relative to the recipient baseline. The same representation supported ASV-level trend forecasting across 24 wastewater treatment plants. Sparse autoencoder analysis resolved MGM2-XLarge token states into a 4,096-feature dictionary spanning taxonomic identity, abundance state, ecological context and technical variation. MGM2 therefore provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery. Highlights[bullet] MGM2 integrates sequence, abundance and community semantics through pretraining on 1.82 million microbiome samples. [bullet]Frozen MGM2 improved macro-AUROC over DeepPhylo by 0.06-0.21 across five temporally held-out MGnify levels. [bullet]MGM2-XLarge reached a response ROC AUC of 0.79 and reduced post-FMT Bray-Curtis distance by 15%. [bullet]A 4,096-feature sparse autoencoder atlas resolves taxonomic, abundance, ecological and technical signals.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine-learning toolbox 95%
- mbImpute: an accurate and robust imputation method for microbiome data 94%
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 94%
Similar papers in this journal
- MaAsLin 3: Refining and extending generalized multivariable linear models for meta-omic association discovery 95%
- mEnrich-seq: Methylation-guided enrichment sequencing of bacterial taxa of interest from microbiome 95%
- Bin Chicken: targeted metagenomic coassembly for the efficient recovery of novel genomes 95%
Similar papers in this journal
Similar papers in this journal
- APOLLO: A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites 95%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 93%
- Spatiotemporal dynamics during niche remodeling by super-colonizing microbiota in the mammalian gut 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.