Controlling the SARS-CoV-2 outbreak, insights from large scale whole genome sequences generated across the world
Phelan, J.; Deelder, W.; Ward, D.; Campino, S.; Hibberd, M. L.; Clark, T. G.
Show abstract
BackgroundSARS-CoV-2 most likely evolved from a bat beta-coronavirus and started infecting humans in December 2019. Since then it has rapidly infected people around the world, with more than 4.5 million confirmed cases by the middle of May 2020. Early genome sequencing of the virus has enabled the development of molecular diagnostics and the commencement of therapy and vaccine development. The analysis of the early sequences showed relatively few evolutionary selection pressures. However, with the rapid worldwide expansion into diverse human populations, significant genetic variations are becoming increasingly likely. The current limitations on social movement between countries also offers the opportunity for these viral variants to become distinct strains with potential implications for diagnostics, therapies and vaccines. MethodsWe used the current sequencing archives (NCBI and GISAID) to investigate 15,487 whole genomes, looking for evidence of strain diversification and selective pressure. ResultsWe used 6,294 SNPs to build a phylogenetic tree of SARS-CoV-2 diversity and noted strong evidence for the existence of two major clades and six sub-clades, unevenly distributed across the world. We also noted that convergent evolution has potentially occurred across several locations in the genome, showing selection pressures, including on the spike glycoprotein where we noted a potentially critical mutation that could affect its binding to the ACE2 receptor. We also report on mutations that could prevent current molecular diagnostics from detecting some of the sub-clades. ConclusionThe worldwide whole genome sequencing effort is revealing the challenge of developing SARS-CoV-2 containment tools suitable for everyone and the need for data to be continually evaluated to ensure accuracy in outbreak estimations.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Emergence and spread of the potential variant of interest (VOI) B.1.1.519 predominantly present in Mexico 97%
- Emergence in Southern France of a new SARS-CoV-2 variant of probably Cameroonian origin harbouring both substitutions N501Y and E484K in the spike protein 97%
- Taxonomic classification methods reveal a new subgenus in the paramyxovirus subfamily Orthoparamyxovirinae 95%
Similar papers in this journal
Similar papers in this journal
- Comprehensive Molecular Epidemiology of Influenza Viruses in Brazil: Insights from a Nationwide Analysis 96%
- Frequent intergenotypic recombination between the non-structural and structural genes is a major driver of epidemiological fitness in caliciviruses 95%
- Diving Deep into Fish Bornaviruses: Uncovering Hidden Diversity and Transcriptional Strategies through Comprehensive Data Mining 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.