Is it the same strain? Defining genomic epidemiology thresholds tailored to individual outbreaks
Duval, A.; opatowski, L.; Brisse, S.
Show abstract
BackgroundEpidemiological surveillance relies on microbial strain typing, which defines genomic relatedness among isolates to identify case clusters and their potential sources. No consensus exists on the choice of thresholds of genomic relatedness to define clusters. While a priori defined thresholds are often applied, outbreak-specific features such as pathogen mutation rate and duration of source contamination should be considered. MethodsWe developed a forward model of bacterial evolution to simulate mutation within a population diversifying at a specific mutation rate, with specific outbreak duration and sample isolation dates. Based on the resulting expected distribution of genetic distances we define a threshold beyond which isolates are considered as not part of the outbreak. We additionally embedded the model into a Markov Chain Monte Carlo inference framework to estimate, from data including sampling dates or isolates genetic variation, the most credible mutation rate or time since source contamination. FindingsA simulation study validated the model over realistic durations and mutation rates. When applied to 16 published datasets describing foodborne outbreaks, our framework consistently identified outliers. Appropriate thresholds for grouping cases were obtained for 14 outbreaks. For the remaining two outbreaks, re-estimation of the duration of outbreak lead to updated threshold values and was more likely, given our model, to result in the observed genetic distances. InterpretationWe propose an evolutionary approach to the single strain conundrum by defining the genetic threshold based on individual outbreak properties. The framework provides an informed estimation of the likelihood of a cluster given the samples epidemiological and microbiological context. This forward model, applicable to foodborne or environmental-source single point case clusters or outbreaks, will be useful for epidemiological surveillance and to guide control measures. FundingThis work was supported financially by the MedVetKlebs project, a component of European Joint Programme One Health EJP, which has received funding from the European Unions Horizon 2020 research and innovation programme under Grant Agreement No 773830. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed for studies published between database inception and April 3, 2021, with the term (threshold OR cut-off OR genetic relatedness) AND (outbreak) AND (cgMLST OR wgMLST OR SNPs) AND (microbial OR bacteria OR bacterial OR pathogen). We found 222 related articles. Most studies define a fixed SNP threshold that relate outbreak strains based on previous observations. One original study identifies outbreak clusters based on transmission events. However, it relies on strong assumptions about molecular clock and transmission processes. Added value of this studyOur study describes a new method based on a forward Wright-Fisher model to find the most credible genetic distance threshold. This method is fast and simple to use with only few assumptions, informed by outbreak duration and pathogen mutation rate. By using SNP or cgMLST pairwise distances and sample collection dates of the outbreak of interest, the algorithm provides context-based guidance to separate outbreak strains from outliers. Implications of all the available evidenceThe fast and easy method developed here enables to move away from a priori defined thresholds. Defining clusters more accurately based on the specific features of outbreaks, and the ability to estimate outbreak duration, will provide the needed precision for epidemiological surveillance and should contribute to leverage molecular epidemiology data more efficiently for the purpose of uncovering contamination sources. Data Availability StatementAll data and code used for this manuscript is available online at https://gitlab.pasteur.fr/BEBP.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Bayesian inference method to estimate transmission trees with multiple introductions; applied to SARS-CoV-2 in Dutch mink farms. 95%
- coiaf: directly estimating complexity of infection with allele frequencies 94%
- Cluster detection with random neighbourhood covering: application to invasive Group A Streptococcal disease 94%
Similar papers in this journal
- Integrated population clustering and genomic epidemiology with PopPIPE 95%
- AB_SA: Tracing the source of bacterial strains based on accessory genes. Application to Salmonella Typhimurium environmental strains 94%
- regentrans: a framework and R package for using genomics to study regional pathogen transmission 94%
Similar papers in this journal
Similar papers in this journal
- Transmission network reconstruction for foot-and-mouth disease outbreaks incorporating farm-level covariates 93%
- Sorting out assortativity: when can we assess the contributions of different population groups to epidemic transmission? 93%
- An accurate and interpretable model for antimicrobial resistance in pathogenic Escherichia coli from livestock and companion animal species 92%
Similar papers in this journal
- Quantifying plasmid movement in drug-resistant Shigella species using phylodynamic inference 92%
- Inferring Mycobacterium bovis transmission between cattle and badgers using isolates from the Randomised Badger Culling Trial 92%
- Amoeba Predation of Cryptococcus: A Quantitative and Population Genomic Evaluation of the Accidental Pathogen Hypothesis 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.