Back

Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy

Mallawaarachchi, S.; Tandon, K.; Rajan, N.; Marcelino, V. R.; Sandhu, S.; Bedoui, S.; Ingle, D. J.; Gunjur, A.; Tonkin-Hill, G.

2026-09-01 microbiology
10.64898/2026.08.30.748153 bioRxiv
Show abstract

Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a species has been used to identify associations between the microbiome and disease. However, this approach overlooks intra-species genetic variation and is susceptible to spurious correlations arising from the compositional nature of abundance data and microbial load. Fast, k-mer-based algorithms can now accurately estimate strain-level Average Nucleotide Identity (ANI) in metagenomes. Despite its value as an orthogonal metric for strain-level analysis, methods for conducting ANI-based association studies remain limited. To address this, we developed StrainSpy, a statistical algorithm that identifies associations between containment ANI and variables of interest across a wide range of study designs, including longitudinal and multi-cohort designs. Re-analysis of a study examining gut microbiota recovery in 12 healthy adults following antibiotic exposure revealed novel strain-level associations, including a reduction in strain-level diversity despite species persistence. Applying StrainSpy to a multi-cohort analysis of 3,414 colorectal cancer metagenomes identified novel strain-level associations with colorectal cancer. However, in a separate collection of microbiome-immunotherapy studies, no individual strain was consistently associated across cohorts. Importantly, across both datasets, StrainSpy informed containment ANI-based machine learning models achieved comparable accuracy to traditional abundance-based methods. StrainSpy is publicly available as an R package github.com/gtonkinhill/strainspy.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.