Strainify: Strain-Level Microbiome Profiling for Low-Coverage Short-Read Metagenomic Datasets
Luo, R. S.; Kille, B.; Vaughan, E. E.; Clark, J. R.; Maresso, A. W.; Nute, M. G.; Treangen, T. J.
Show abstract
MotivationStrain-level microbiome profiling has revealed key insights into microbial community composition and strain dynamics. However, accurate strain-level analysis remains challenging due to limited linkage information, ambiguous read mapping, and complicating factors such as genome similarity, sequencing depth, and community complexity. These challenges are especially pronounced for short-read metagenomic data when estimating the relative abundances of multiple strains, a task critical for genotype-phenotype association studies. ResultsTo address this gap, we present Strainify, which enables accurate strain-level abundance estimation from short-read metagenomes with as little as 1% genome coverage. Specifically, Strainify combines (1) identification of informative variants via core genome alignment, (2) filtering of confounding variants via a window-based test, and (3) maximum likelihood estimation of strain abundances. A Shannon entropy-weighted version of the model further improves robustness in noisy, low-coverage settings by downweighting sites with low information content. Across simulated communities of varying complexity, Strainify consistently outperformed existing approaches. On mock community sequencing data, Strainifys estimates aligned more closely with reference abundances. When applied to a longitudinal gut microbiome dataset, Strainify successfully recapitulated the reported temporal dynamics of Bacteroides ovatus strain groups, demonstrating its ability to recover biologically meaningful patterns from real-world metagenomes. Together, these results establish Strainify as a robust and versatile solution for accurate strain-level abundance estimation in short-read, low-coverage microbiome studies. AvailabilityThe Strainify code and results are available at: https://github.com/treangenlab/Strainify
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MDITRE: scalable and interpretable machine learning for predicting host status from temporal microbiome dynamics 96%
- BiG-MAP: an automated pipeline to profile metabolic gene cluster abundance and expression in microbiomes 95%
- Illumina Complete Long Read Assay yields contiguous bacterial genomes from human gut metagenomes 95%
Similar papers in this journal
- Identifying unmeasured heterogeneity in microbiome data via quantile thresholding (QuanT) 97%
- BinaRena: a dedicated interactive platform for human-guided exploration and binning of metagenomes 97%
- MetaPro: A scalable and reproducible data processing and analysis pipeline for metatranscriptomic investigation of microbial communities 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.