Back

Diagnostic Potential of Shallow Depth Gut Metagenomics Sequencing for Atherosclerotic Cardiovascular Disease Risk Stratification

Sing, J. C.; Whitley, O.; Davis, S.

2023-10-04 bioinformatics
10.1101/2023.10.02.560614 bioRxiv
Show abstract

BackgroundCardiovascular disease, specifically atherosclerotic cardiovascular disease (ACVD), presents a burden on society in terms of financial resources and quality of life that is only expected to grow in the coming years. Current diagnostic methods lack the sensitivity or specificity to confidently screen for atherosclerotic cardiovascular disease and are either costly, invasive, or both. Building on previous research linking ACVD to changes in the microbiome, we hypothesized that one could build a machine learning classifier for ACVD that makes use of shallow depth microbiome data. Methods and FindingsRaw metagenomics Illumina paired end sequencing data was downloaded from the European Bioinformatics Institute (EBI) public database under accession ERP023788. The dataset includes 383 Han Chinese subjects of which 170 are control subjects and 214 are ACVD subjects. The raw sequencing data was subsampled using seqtk to generate in silico shallow depth datasets at various depths. Each subsampled experiment was processed through KneadData for quality control, and finally relative abundances were acquired through Kraken2. Based on in-silico down sampling experiments, we demonstrate that species level information is still captured at read depths as low as 50K reads per sample, consistent with current literature. In addition, the shallow depth data (50K) contains relevant taxa information for stratification of ACVD patients. Differential expression analysis identified a variety of Streptococcus species and Escherichia coli enriched in ACVD patients and a few low abundance Bacteroides species in ACVD patients, which have been previously reported. ConclusionHere, using publicly available data, we show with in silico experiments that species-level microbiome information is preserved at low sequencing depths that would allow a screening or diagnostic tool to be cost-competitive with currently used methods such as stress ECGs and stress ECHO tests. Additionally, we show that we can indeed make a microbiome-based model with performance metrics comparable to or better than front-line or mid-tier tests.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.