Back

Annotation-free sequence-level multimodal graph learning reveals gut microbiome signatures of atherosclerotic cardiovascular disease

Hu, W.; Wang, W.; Zhang, L.; Fu, Y. V.

2026-05-27 bioinformatics
10.64898/2026.05.24.727482 bioRxiv
Show abstract

BackgroundAtherosclerotic cardiovascular disease (ACVD) remains one of the leading causes of mortality, and the gut microbiome has been implicated in ACVD-related metabolic and inflammatory processes. However, current microbiome-based disease prediction frameworks predominantly rely on annotation-dependent taxonomic or pathway-level abundance profiles. While effective for community-level association analysis, these approaches compress metagenomic information into predefined biological categories and may consequently obscure fine-grained sequence variation, uncharacterized microbial fragments, regulatory sequence signals and sequence-based structure information embedded within metagenomic DNA. Whether disease-relevant microbiome information can be directly learned from raw metagenomic sequences without prior annotation remains incompletely explored. ResultsWe developed MLMGCN-CVD, an annotation-free sequence-level multimodal graph-learning framework for ACVD prediction directly from gut metagenomic DNA fragments. Rather than relying on taxonomic aggregation, MLMGCN-CVD integrates pretrained genomic language-model embeddings, sequence-derived structural priors, and read-mapping quantitative evidence to learn disease-associated representations at the DNA fragment level. In a primary cohort comprising 218 ACVD patients and 187 healthy controls, MLMGCN-CVD achieved an area under the receiver operating characteristic curve (AUC) of 0.978 under grouped 10-fold cross-validation. External validation on an independent cohort (PRJNA615842) yielded an AUC of 0.940, outperforming relative- abundance-, pathway-, and sequence-only reference frameworks. Importantly, model-prioritized fragments independently recovered multiple microbiome-associated functions previously implicated in ACVD, including trimethylamine (TMA) metabolism, O-antigen of lipopolysaccharide (LPS) biosynthesis, phosphotransferase systems (PTS), and Enterobacteriaceae-associated modules. Beyond these established signatures, Rfam-supported analyses identified candidate regulatory RNA-associated signals, including riboswitches and Bacterial small RNAs, suggesting that regulatory sequence information may contribute previously underexplored discriminatory signals within the ACVD gut microbiome. Several representative sequence elements further showed significant correlations with host cardiometabolic indicators. ConclusionsOur findings suggest that microbiome-associated disease signals extend beyond conventional taxonomic or pathway abundance summaries and can be directly inferred from raw metagenomic sequences. By shifting microbiome modelling from annotation-dependent community aggregation toward annotation-free sequence-level representation learning, MLMGCN-CVD provides a framework for uncovering biologically informative regulatory and functional signals embedded within metagenomic data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.