Characterization of the substitution hotspots in SARS-CoV-2 genome using BioAider and detection of a SR-rich region in N protein providing further evidence of its animal origin
Zhou, Z.-J.; Qiu, Y.; Ge, X.-Y.
Show abstract
The novel human coronavirus (SARS-CoV-2) causes the coronavirus disease 2019 (COVID-19) pandemic worldwide. The increasing sequencing data have shown abundant single nucleotide variations in SARS-CoV-2 genome. However, it is difficult to quickly analyze genomic variation and screen key mutations of SARS-CoV-2. In this study, we developed a visual program, named BioAider, for quick and convenient sequence annotation and mutation analysis on multiple genome-sequencing data. Using BioAider, we conducted a comprehensive genome variation analysis on 3,240 sequences of SARS-CoV-2 genome. Herein, we detected 14 substitution hotspots within SARS-CoV-2 genome, including 10 non-synonymous and 4 synonymous ones. Among these hotspots, NSP13-Y541C was predicted to be a crucial substitution which might affect the unwinding activity of NSP13, a key protein for viral replication. Besides, we also found 3 groups of potentially linked substitution hotspots which were worth further study. In particular, we discovered a SR-rich region (aa 184-204) on the N protein of SARS-CoV-2 distinct from SARS-CoV, indicating more complex replication mechanism and unique N-M interaction of SARS-CoV-2. Interestingly, the quantity of SRXX repeat fragments in the SR-rich region well reflected the evolutionary relationship among SARS-CoV-2 and SARS-CoV-2 related animal coronaviruses, providing further evidence of its animal origin. Overall, we developed an efficient tool for rapid identification of mutations, identified substitution hotspots in SARS-CoV-2 genomes, and detected a distinctive polymorphism SR-rich region in N protein. This tool and the detected hotspots could facilitate the viral genomic study and may contribute for screening antiviral target sites.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Whole Genome Sequencing Analysis of Spike D614G Mutation Reveals Unique SARS-CoV-2 Lineages of B.1.524 and AU.2 in Malaysia 97%
- SARS-CoV-2: Proof of recombination between strains and emergence of possibly more virulent ones 97%
- Analysis the molecular similarity of least common amino acid sites in ACE2 receptor to predict the potential susceptible species for SARS-CoV-2 96%
Similar papers in this journal
- Whole Genome Comparison of Pakistani Corona Virus with Chinese and US Strains along with its Predictive Severity of COVID-19 96%
- SARS-CoV-2 mutations altering regulatory properties: deciphering host's and virus's perspectives 96%
- Analysis of single nucleotide polymorphisms between 2019-nCoV genomes and its impact on codon usage 96%
Similar papers in this journal
- Comparative Analysis of Human Coronaviruses Focusing on Nucleotide Variability and Synonymous Codon Usage Pattern 97%
- SARS-CoV-2 transcriptome analysis and molecular cataloguing of immunodominant epitopes for multi-epitope based vaccine design 95%
- A Computational Approach to Design Potential siRNA Molecules as a Prospective Tool for Silencing Nucleocapsid Phosphoprotein and Surface Glycoprotein Gene of SARS-CoV-2 94%
Similar papers in this journal
- Deciphering inhibitory mechanism of coronavirus replication through host miRNAs-RNA-dependent RNA polymerase (RdRp) interactome 96%
- How the replication and transcription complex functions in jumping transcription of SARS-CoV-2 96%
- Genomic variations in SARS-CoV-2 genomes from Gujarat: Underlying role of variants in disease epidemiology 95%
Similar papers in this journal
- Host and infectivity prediction of Wuhan 2019 novel coronavirus using deep learning algorithm 97%
- Emergence of SARS-CoV-2 Omicron Variant JN.1 in Tamil Nadu, India - Clinical Characteristics and Novel Mutations 95%
- Machine learning prediction of antiviral-HPV protein interactions for anti-HPV pharmacotherapy 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.