A new method for detecting mixed Mycobacterium tuberculosis infection and reconstructing constituent strains provides insights into transmission
Sobkowiak, B.; Cudahy, P.; Chitwood, M. H.; Clark, T. G.; Colijn, C.; Grandjean, L.; Walter, K. S.; Crudu, V.; Cohen, T.
Show abstract
BackgroundMixed infection with multiple strains of the same pathogen in a single host can present clinical and analytical challenges. Whole genome sequence (WGS) data can identify signals of multiple strains in samples, though the precision of previous methods can be improved. Here, we present MixInfect2, a new tool to accurately detect mixed samples from Mycobacterium tuberculosis WGS data. We then evaluate three approaches for reconstructing the underlying mixed constituent strain sequences. This allows these samples to be included in downstream analysis to gain insights into the epidemiology and transmission of mixed infections. MethodsWe employed a Gaussian mixture model to cluster allele frequencies at mixed sites (hSNPs) in each sample to identify signals of multiple strains. Building upon our previous tool, MixInfect, we increased the accuracy of classifying in vitro mixed samples through multiple improvements to the bioinformatic pipeline. Major and minor proportion constituent strains were reconstructed using three approaches and assessed by comparing the estimated sequence to the known constituent strain sequence. Lastly, mixed infections in a real-world Mycobacterium tuberculosis population from Moldova were detected with MixInfect2 and clusters of recent transmission that included major and minor constituent strains were built. ResultsAll 36/36 in vitro mixed and 12/12 non-mixed samples were correctly classified with MixInfect2, and major strain proportions estimated with high accuracy, outperforming previous tools. Reconstructed major strain sequences closely matched the true constituent sequence by taking the allele at the highest frequency at hSNPs, while the best performing approach to reconstruct the minor proportion strain sequence was identifying the closest non-mixed isolate in the same population, though no approach was effective when the minor strain proportion was at 5%. Finally, fewer mixed infections were identified in Moldova than previous estimates (6.6% vs 17.4%) and we found multiple instances where the constituent strains of mixed samples were present in transmission clusters. ConclusionsMixInfect2 accurately detects samples with evidence of mixed infection from WGS data and provides an excellent estimate of the mixture proportions. While there are limitations in reconstructing the constituent strain sequences of mixed samples, we present recommendations for the best approach to include these isolates in further analyses.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SNPPar: identifying convergent evolution and other homoplasies from microbial whole-genome alignments 95%
- Unravelling the population structure and transmission patterns of Mycobacterium tuberculosis in Mozambique, a high TB/HIV burden country 95%
- Tracking SARS-CoV-2 variants of concern in wastewater: an assessment of nine computational tools using simulated genomic data 94%
Similar papers in this journal
Similar papers in this journal
- A needle in a haystack: metagenomic DNA sequencing to quantify Mycobacterium tuberculosis DNA and diagnose tuberculosis 94%
- Sensitive and modular amplicon sequencing of Plasmodium falciparum diversity and resistance for research and public health 93%
- An Escherichia coli ST131 pangenome atlas reveals population structure and evolution across 4,071 isolates 92%
Similar papers in this journal
- Deciphering the tangible spatio-temporal spread of a 25 years tuberculosis outbreak boosted by social determinants 93%
- Development of an amplicon nanopore sequencing strategy for detection of mutations conferring intermediate resistance to vancomycin in Staphylococcus aureus strains 92%
- Targeted Hybridization Capture of SARS-CoV-2 and Metagenomics Enables Genetic Variant Discovery and Nasal Microbiome Insights 92%
Similar papers in this journal
- Hash-based core genome multi-locus sequencing typing for Clostridium difficile 94%
- Accurate and Reproducible Whole-Genome Genotyping for Bacterial Genomic Surveillance with Nanopore Sequencing Data 93%
- Identifying indels from WGS short reads of haploid genomes distinguishes variant-calling algorithms 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.