Back

Genomics-informed outbreak investigations of SARS-CoV-2 using civet

O'Toole, A. N.; Hill, V.; Jackson, B.; Dewar, R.; Sahadeo, N.; Colquhoun, R.; Rooke, S.; McCrone, J. T.; McHugh, M.; Nicholls, S.; Poplawski, R.; The COVID-19 Genomics UK (COG-UK) Consortium, ; COVID-19 Impact Project (Trinidad & Tobago Group), ; Aanensen, D.; Holden, M.; Connor, T. R.; Loman, N.; Goodfellow, I. G.; Carrington, C.; Templeton, K.; Rambaut, A.

2021-12-14 epidemiology
10.1101/2021.12.13.21267267 medRxiv
Show abstract

The scale of data produced during the SARS-CoV-2 pandemic has been unprecedented, with more than 5 million sequences shared publicly at the time of writing. This wealth of sequence data provides important context for interpreting local outbreaks. However, placing sequences of interest into national and international context is difficult given the size of the global dataset. Often outbreak investigations and genomic surveillance efforts require running similar analyses again and again on the latest dataset and producing reports. We developed civet (cluster investigation and virus epidemiology tool) to aid these routine analyses and facilitate virus outbreak investigation and surveillance. Civet can place sequences of interest in the local context of background diversity, resolving the query into different catchments and presenting the phylogenetic results alongside metadata in an interactive, distributable report. Civet can be used on a fine scale for clinical outbreak investigation, for local surveillance and cluster discovery, and to routinely summarise the virus diversity circulating on a national level. Civet reports have helped researchers and public health bodies feedback genomic information in the appropriate context within a timeframe that is useful for public health.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.