SARS-CoV-2's evolutionary capacity is mostly driven by host antiviral molecules
Lamb, K. D.; Luka, M. M.; Saathoff, M.; Orton, R.; Phan, M.; Cotten, M.; Yuan, K.; Robertson, D. L.
Show abstract
The COVID-19 pandemic has been characterised by sequential variant-specific waves shaped by viral, individual human and population factors. SARS-CoV-2 variants are defined by their unique combinations of mutations and there has been a clear adaptation to human infection since its emergence in 2019. Here we use machine learning models to identify shared signatures, i.e., common underlying mutational processes, and link these to the subset of mutations that define the variants of concern (VOCs). First, we examined the global SARS-CoV-2 genomes and associated metadata to determine how viral properties and public health measures have influenced the magnitude of waves, as measured by the number of infection cases, in different geographic locations using regression models. This analysis showed that, as expected, both public health measures and not virus properties alone are associated with the rise and fall of regional SARS-CoV-2 reported infection numbers. This impact varies geographically. We attribute this to intrinsic differences such as vaccine coverage, testing and sequencing capacity, and the effectiveness of government stringency. In terms of underlying evolutionary change, we used non-negative matrix factorisation to observe three distinct mutational signatures, unique in their substitution patterns and exposures from the SARS-CoV-2 genomes. Signatures 0, 1 and 3 were biased to C[->]T, T[->]C/A[->]G and G[->]T point mutations as would be expected of host antiviral molecules APOBEC, ADAR and ROS effects, respectively. We also observe a shift amidst the pandemic in relative mutational signature activity from predominantly APOBEC-like changes to an increasingly high proportion of changes consistent with ADAR editing. This could represent changes in how the virus and the host immune response interact, and indicates how SARS-CoV-2 may continue to accumulate mutations in the future. Linkage of the detected mutational signatures to the VOC defining amino acids substitutions indicates the majority of SARS-CoV-2s evolutionary capacity is likely to be associated with the action of host antiviral molecules rather than virus replication errors.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Intrahost SARS-CoV-2 k-mer identification method (iSKIM) for rapid detection of mutations of concern reveals emergence of global mutation patterns 96%
- Hacking the diversity of SARS-CoV-2 and SARS-like coronaviruses in human, bat and pangolin populations 95%
- Origins and Evolution of Seasonal Human Coronaviruses 95%
Similar papers in this journal
- Structural impact of synonymous mutations in six SARS-CoV-2 Variants of Concern 96%
- Machine learning using intrinsic genomic signatures for rapid classification of novel pathogens: COVID-19 case study 95%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 95%
Similar papers in this journal
- Nanopore and Illumina Sequencing Reveal Different Viral Populations from Human Gut Samples 94%
- Phylogenomics and population genomics of SARS-CoV-2 in Mexico reveals variants of interest (VOI) and a mutation in the Nucleocapsid protein associated with symptomatic versus asymptomatic carriers 94%
- Genomic Epidemiology of SARS-CoV-2 in Norfolk, UK, March 2020 - December 2022 93%
Similar papers in this journal
- SARS-CoV-2 within-host and in-vitro genomic variability and sub-genomic RNA levels indicate differences in viral expression between clinical and in-vitro cohorts. 95%
- In depth characterization of an archaeal virus-host system reveals numerous virus exclusion mechanisms 95%
- Global Geographic and Temporal Analysis of SARS-CoV-2 Haplotypes Normalized by COVID-19 Cases during the Pandemic 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.