Annotation-free sequence-level multimodal graph learning reveals gut microbiome signatures of atherosclerotic cardiovascular disease
Hu, W.; Wang, W.; Zhang, L.; Fu, Y. V.
Show abstract
BackgroundAtherosclerotic cardiovascular disease (ACVD) remains one of the leading causes of mortality, and the gut microbiome has been implicated in ACVD-related metabolic and inflammatory processes. However, current microbiome-based disease prediction frameworks predominantly rely on annotation-dependent taxonomic or pathway-level abundance profiles. While effective for community-level association analysis, these approaches compress metagenomic information into predefined biological categories and may consequently obscure fine-grained sequence variation, uncharacterized microbial fragments, regulatory sequence signals and sequence-based structure information embedded within metagenomic DNA. Whether disease-relevant microbiome information can be directly learned from raw metagenomic sequences without prior annotation remains incompletely explored. ResultsWe developed MLMGCN-CVD, an annotation-free sequence-level multimodal graph-learning framework for ACVD prediction directly from gut metagenomic DNA fragments. Rather than relying on taxonomic aggregation, MLMGCN-CVD integrates pretrained genomic language-model embeddings, sequence-derived structural priors, and read-mapping quantitative evidence to learn disease-associated representations at the DNA fragment level. In a primary cohort comprising 218 ACVD patients and 187 healthy controls, MLMGCN-CVD achieved an area under the receiver operating characteristic curve (AUC) of 0.978 under grouped 10-fold cross-validation. External validation on an independent cohort (PRJNA615842) yielded an AUC of 0.940, outperforming relative- abundance-, pathway-, and sequence-only reference frameworks. Importantly, model-prioritized fragments independently recovered multiple microbiome-associated functions previously implicated in ACVD, including trimethylamine (TMA) metabolism, O-antigen of lipopolysaccharide (LPS) biosynthesis, phosphotransferase systems (PTS), and Enterobacteriaceae-associated modules. Beyond these established signatures, Rfam-supported analyses identified candidate regulatory RNA-associated signals, including riboswitches and Bacterial small RNAs, suggesting that regulatory sequence information may contribute previously underexplored discriminatory signals within the ACVD gut microbiome. Several representative sequence elements further showed significant correlations with host cardiometabolic indicators. ConclusionsOur findings suggest that microbiome-associated disease signals extend beyond conventional taxonomic or pathway abundance summaries and can be directly inferred from raw metagenomic sequences. By shifting microbiome modelling from annotation-dependent community aggregation toward annotation-free sequence-level representation learning, MLMGCN-CVD provides a framework for uncovering biologically informative regulatory and functional signals embedded within metagenomic data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 95%
- uniPort: a unified computational framework for single-cell data integration with optimal transport 95%
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 95%
Similar papers in this journal
- Microbial general model: Leveraging large language model for contextualized microbiome analysis 96%
- Deciphering 3'UTR mediated gene regulation using interpretable deep representation learning 95%
- Long-read sequencing reveals extensive DNA methylations in human gut phagenome contributed by prevalently phage-encoded methyltransferases 94%
Similar papers in this journal
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 94%
- StereoMM: A Graph Fusion Model for Integrating Spatial Transcriptomic Data and Pathological Images 94%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 93%
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 94%
- ModiDeC: a multi-RNA modification classifier for direct nanopore sequencing 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.