Back

Three adjacent nucleotide changes spanning two residues in SARS-CoV-2 nucleoprotein: possible homologous recombination from the transcription-regulating sequence

Leary, S.; Gaudieri, S.; Chopra, A.; Pakala, S.; Alves, E.; John, M.; Das, S.; Mallal, S.; Phillips, P. J.

2020-04-11 microbiology
10.1101/2020.04.10.029454 bioRxiv
Show abstract

BackgroundGenetic variations across the SARS-CoV-2 genome may influence transmissibility of the virus and the hosts anti-viral immune response, in turn affecting the frequency of variants over-time. In this study, we examined the adjacent amino acid polymorphisms in the nucleocapsid (R203K/G204R) of SARS-CoV-2 that arose on the background of the spike D614G change and describe how strains harboring these changes became dominant circulating strains globally. MethodsDeep sequencing data of SARS-CoV-2 from public databases and from clinical samples were analyzed to identify and map genetic variants and sub-genomic RNA transcripts across the genome. ResultsSequence analysis suggests that the three adjacent nucleotide changes that result in the K203/R204 variant have arisen by homologous recombination from the core sequence (CS) of the leader transcription-regulating sequence (TRS) rather than by stepwise mutation. The resulting sequence changes generate a novel sub-genomic RNA transcript for the C-terminal dimerization domain of nucleocapsid. Deep sequencing data from 981 clinical samples confirmed the presence of the novel TRS-CS-dimerization domain RNA in individuals with the K203/R204 variant. Quantification of sub-genomic RNA indicates that viruses with the K203/R204 variant may also have increased expression of sub-genomic RNA from other open reading frames. ConclusionsThe finding that homologous recombination from the TRS may have occurred since the introduction of SARS-CoV-2 in humans resulting in both coding changes and novel sub-genomic RNA transcripts suggests this as a mechanism for diversification and adaptation within its new host.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.