Back

Infection, Genetics and Evolution

Elsevier BV

All preprints, ranked by how well they match Infection, Genetics and Evolution's content profile, based on 42 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
The new Coronavirus (SARS-CoV-2) in Central America: Demographic-spatial simulations, Analyses of Molecular Variance (AMOVA) and Neutrality Tests in complete genomes from Belize, Guatemala, Cuba, Jamaica and Puerto Rico

Ramos, R. d. S.; Venancio, D. B. R.; Da Silva, E. D. A. B.; de Albuquerque, R. M.; Felix, P. T.

2021-01-02 infectious diseases 10.1101/2020.12.26.20248872 medRxiv
Top 0.1%
35.3%
Show abstract

In this work, we evaluated the levels of genetic diversity in 38 complete genomes of SARS-CoV-2 from five Central American countries (Belize, Guatemala, Cuba, Jamaica and Puerto Rico) with 04, 10, 2, 8 and 14 haplotypes, respectively, with an extension of up to 29,885 bp. All sequences were publicly available on the National Biotechnology Information Center (NCBI) platform. Using specific methodologies for paired FST, AMOVA, mismatch, demographic-spatial expansion, molecular diversity and for the time of evolutionary divergence, it was possible to notice that only 79 sites remained conserved and that the high number of polymorphisms found helped to establish a clear pattern of genetic non-structuring, based on the time of divergence between the groups. The analyses also showed that significant evolutionary divergences within and between the five countries corroborate the fact that possible rapid and silent mutations are responsible for the increase in genetic variability of the Virus, a fact that would hinder the work with molecular targets for vaccines and medications in general.

2
In Silico Epitope Prediction And VP1 Modelling For Foot-And-Mouth Serotype SAT 2 For Vaccine Design In East Africa

Udahemuka, J. C.; Lunayo, A.; Obiero, G. O.; Aboge, G. O.; Lebea, P. J.

2021-09-15 immunology 10.1101/2021.09.13.460008 medRxiv
Top 0.1%
31.2%
Show abstract

Foot and Mouth Disease Virus has seven distinct, geographically localized, serotypes and a vaccination targeting one serotype does not confer immunity against another serotype. The use of inactivated vaccines is not safe and confers an immunity with a relatively shorter time. Using the VP1 sequences isolated in East Africa, we have predicted epitopes able to induce humoral and cell-mediated immunity in cattle. The Wu-Kabat variability index calculated in this study reflects the variable, including the known GH loop, and conserved regions, with the latter being good candidates for region-tailored vaccine design. Furthermore, we modelled the identified epitopes on a 3D model (PDB ID:5aca) to represent the epitopes structurally. This study can be used for in vitro and in vivo experiments.

3
Evolutionary analysis of SARS-CoV-2 spike protein for its different clades

Pereson, M. J.; Flichman, D.; Martinez, A.; Bare, P.; Garcia, G.; Di Lello, F.

2020-11-25 evolutionary biology 10.1101/2020.11.24.396671 medRxiv
Top 0.1%
31.1%
Show abstract

ObjectiveThe spike protein of SARS-CoV-2 has become the main target for antiviral and vaccine development. Despite its relevance, there is scarce information about its evolutionary traces. The aim of this study was to investigate the diversification patterns of the spike for each clade of SARS-CoV-2 through different approaches. MethodsTwo thousand and one hundred sequences representing the seven clades of the SARS-CoV-2 were included. Patterns of genetic diversifications and nucleotide evolutionary rate were estimated for the spike genomic region. ResultsThe haplotype networks showed a star shape, where multiple haplotypes with few nucleotide differences diverge from a common ancestor. Four hundred seventy nine different haplotypes were defined in the seven analyzed clades. The main haplotype, named Hap-1, was the most frequent for clades G (54%), GH (54%), and GR (56%) and a different haplotype (named Hap-252) was the most important for clades L (63.3%), O (39.7%), S (51.7%), and V (70%). The evolutionary rate for the spike protein was estimated as 1.08 x 10-3 nucleotide substitutions/site/year. Moreover, the nucleotide evolutionary rate after nine months of pandemic was similar for each clade. ConclusionsIn conclusion, the present evolutionary analysis is relevant since the spike protein of SARS-CoV-2 is the target for most therapeutic candidates; besides, changes in this protein could have consequences on viral transmission, response to antivirals and efficacy of vaccines. Moreover, the evolutionary characterization of clades improves knowledge of SARS-CoV-2 and deserves to be assessed in more detail since re-infection by different phylogenetic clades has been reported.

4
Levels of genetic diversity of SARS-CoV-2 virus: reducing speculations about the genetic variability of the virus in South America

Filho, C. B. d. N.; Ramos, R. d. S.; Paulino, A. J.; Venancio, D. B. R.; Felix, P. T.

2020-09-15 genetics 10.1101/2020.09.14.296491 medRxiv
Top 0.1%
27.4%
Show abstract

In this work, we evaluated the levels of genetic diversity in 38 complete Genomes of SARS-CoV-2 from six countries in South America, using specific methodologies for paired FST, AMOVA, mismatch, demographic and spatial expansions, molecular diversity and for the time of evolutionary divergence. The analyses showed non-significant evolutionary divergences within and between the six countries, as well as a significant similarity to the time of genetic evolutionary divergence between all populations. Thus, it seems safe to affirm that we will find similar results for the other Countries of South America, reducing speculation about the existence of rapid and silent mutations that, although there are as we have shown in this work, do not increase, until this moment, the genetic variability of the Virus, a fact that would hinder the work with molecular targets for vaccines and drugs in general.

5
Evolution of tandem repeats in putative CSP to enhance its function: A recent and exclusive event in Plasmodium vivax in India

Dash, M.; Pande, V.; DAS, A.; Sinha, A.

2023-11-28 evolutionary biology Community evaluation 10.1101/2023.11.28.568961 medRxiv
Top 0.1%
26.9%
Show abstract

The molecular hitchhiking model proposes that linked non-coding regions also undergo fixation, while fixing a beneficial allele in a population. This concept can be applied to identify loci with functional and evolutionary significance. Putative circumsporozoite protein (CSP) in Plasmodium vivax (PvpuCSP) identified following the molecular hitchhiking model, holds evolutionary significance. We investigated the extent of genetic polymorphism in PvpuCSP and the role of natural selection which shapes the genetic composition and maintains the diversity in P. vivax isolates from India. Sequencing the putative CSP of P. vivax (PvpuCSP) in 71 isolates revealed a well-conserved N- and C-terminal, constituting around 80% of the gene. PCR amplification and sequencing validated extensive diversity in the repeat region, ranging from 1.8 to 2.2 kb towards the C-terminal, identifying 37 different alleles from 71 samples. The recent and exclusive accumulation of repeats in puCSP within P. vivax highlights its highly variable length polymorphism, making it a potential marker for estimating diversity and infection complexity. Episodic diversifying selection in the PvpuCSP repeat region, evidenced by statistically significant p-values and likelihood ratios, enhances amino acid diversity at various phylogenetic levels, facilitating adaptation for accommodating different substrates for degradation.

6
Bioinformatic analysis of shared B and T cell epitopes amongst relevant coronaviruses to human health: Is there cross-protection?

Pacheco-Olvera, D. L.; Saint Remy-Hernandez, S.; Acevedo-Ochoa, E.; Arriaga-Pizano, L.; Cerbulo-Vazquez, A.; Ferat-Osorio, E.; Rivera-Hernandez, T.; Lopez-Macias, C.

2020-07-15 immunology 10.1101/2020.07.14.202887 medRxiv
Top 0.1%
22.8%
Show abstract

Within the last 30 years 3 coronaviruses, SARS-CoV, MERS-CoV and SARS-CoV-2, have evolved and adapted to cause disease and spread amongst the human population. From the three, SARS-CoV-2 has spread world-wide and to July 2020 it has been responsible for more than 11 million confirmed cases and over half a million deaths. In the absence of an effective treatment or vaccine, social distancing has been the most effective measure to control the pandemic. However it has become evident that as the virus spreads the only tool that will allow us to fully control it is an effective vaccine. There are currently more than 150 vaccine candidates in different stages of development using a variety of viral antigens, with the S protein being the most targeted antigen. Although some new experimental evidence suggests cross-reacting responses between coronaviruses are present in the population, it remains unknown whether potential shared antigens between different coronaviruses could provide cross-protection. Given that coronaviruses are emerging pathogens and continue to represent a threat to global health, the development of a SARS-Cov-2 vaccine that could provide universal protection against other coronaviruses should be pushed forward. Here we present a thorough review of reported B and T cell epitopes shared between SARS-CoV-2 and other relevant coronaviruses, in addition we used web-based tools to predict novel B and T cell epitopes that have not been reported before. Analysis of experimental evidence that is constantly emerging complemented with the findings of this study allow us support the hypothesis that cross-reactive responses, particularly those coming from T cells, might play a key role in controlling infection by SARS-CoV-2. We hope that with the evidence presented in this manuscript we provide arguments to encourage the study of cross-reactive responses in order to elucidate their role in immunity to SARS-CoV-2. Finally we expect our findings will aid targeted analysis of antigen-specific immune responses and guide future vaccine design aiming to develop a cross reactive effective vaccine against respiratory diseases caused by coronaviruses.

7
Selective pressure on membrane proteins drives the evolution of Helicobacter pylori Colombian subpopulations

Guevara Tique, A. A.; Torres, R. C.; Castro Valencia, F. L.; Suarez, J. J.; Criollo Rayo, A. A.; Bravo, M. M.; Carvajal Carmona, L. G.; Echeverry de Polanco, M. M.; Bohorquez Lozano, M. E.; Torres, J.

2021-12-16 genomics 10.1101/2021.12.14.472690 medRxiv
Top 0.1%
22.2%
Show abstract

Helicobacter pylori have coevolved with mankind since its origins, adapting to different human groups. In America H. pylori has evolved in several subpopulations specific for regions or even countries. In this study we analyzed the genome of 163 Colombian strains along with 1,113 strains that represent worldwide H. pylori populations to better discern the ancestry and adaption to Colombian people. Population structure was inferred with FineStructure and chromosome painting identifying the proportion of ancestries in Colombian isolates. Phylogenetic relationship was analyzed using the SNPs present in the core genome. Also, a Fst analysis was done to identify the gene variants with the strongest fixation in the identified Colombian subpopulations in relation to their parent population hspSWEurope. Worldwide, population structure analysis allowed the identification of two Colombian subpopulations, the previously described hspSWEuropeColombia and a novel subpopulation named hspColombia. In addition, three subgroups of H. pylori were identified within hspColombia that follow their geographic origin. The Colombian H. pylori subpopulations represent an admixture of European, African and Native indigenous ancestry; although some genomes showed a high proportion of self-identity, suggesting a strong adaption to these mestizo Colombian groups. The Fst analysis identified 82 SNPs significantly fixed in 26 genes of the hspColombia subpopulation that encode mainly for outer membrane proteins and proteins involved in central metabolism. The strongest fixation indices were identified in genes encoding the membrane proteins HofC, HopE, FrpB-4 and Sialidase A. These findings demonstrate that H. pylori has evolved in Colombia to give rise to subpopulations following a geographical structure, evolving to an autochthonous genetic pool, drive by a positive selective pressure especially on genes encoding for outer membrane proteins.

8
Analysis of the mutation dynamics of SARS-CoV-2 reveals the spread history and emergence of RBD mutant with lower ACE2 binding affinity

JIA, Y.; Shen, G.; Zhang, Y.; Huang, K.-S.; Ho, H.-Y.; Hor, W.-S.; Yang, C.-H.; Li, C.; Wang, W.-L.

2020-04-11 evolutionary biology 10.1101/2020.04.09.034942 medRxiv
Top 0.1%
19.0%
Show abstract

Monitoring the mutation dynamics of SARS-CoV-2 is critical for the development of effective approaches to contain the pathogen. By analyzing 106 SARS-CoV-2 and 39 SARS genome sequences, we provided direct genetic evidence that SARS-CoV-2 has a much lower mutation rate than SARS. Minimum Evolution phylogeny analysis revealed the putative original status of SARS-CoV-2 and the early-stage spread history. The discrepant phylogenies for the spike protein and its receptor binding domain proved a previously reported structural rearrangement prior to the emergence of SARS-CoV-2. Despite that we found the spike glycoprotein of SARS-CoV-2 is particularly more conserved, we identified a receptor binding domain mutation that leads to weaker ACE2 binding capability based on in silico simulation, which concerns a SARS-CoV-2 sample collected on 27th January 2020 from India. This represents the first report of a significant SARS-CoV-2 mutant, and requires attention from researchers working on vaccine development around the world. HighlightsO_LIBased on the currently available genome sequence data, we provided direct genetic evidence that the SARS-COV-2 genome has a much lower mutation rate and genetic diversity than SARS during the 2002-2003 outbreak. C_LIO_LIThe spike (S) protein encoding gene of SARS-COV-2 is found relatively more conserved than other protein-encoding genes, which is a good indication for the ongoing antiviral drug and vaccine development. C_LIO_LIMinimum Evolution phylogeny analysis revealed the putative original status of SARS-CoV-2 and the early-stage spread history. C_LIO_LIWe confirmed a previously reported rearrangement in the S protein arrangement of SARS-COV-2, and propose that this rearrangement should have occurred between human SARS-CoV and a bat SARS-CoV, at a time point much earlier before SARS-COV-2 transmission to human. C_LIO_LIWe provided first evidence that a mutated SARS-COV-2 with reduced human ACE2 receptor binding affinity have emerged in India based on a sample collected on 27th January 2020. C_LI

9
In silico assessment of immune cross protection between BCoV and SARS-CoV-2

Beirao, B. C. B.; Querne, L. B.; Zettel Bastos, F.; Adur, M. d. A.; Cavalheiro, V. L.

2024-01-26 immunology 10.1101/2024.01.25.577193 medRxiv
Top 0.1%
19.0%
Show abstract

BackgroundHumans have long shared infectious agents with cattle, and the bovine-derived human common cold OC-43 CoV is a not-so-distant example of cross-species viral spill over of coronaviruses. Human exposure to the Bovine Coronavirus (BCoV) is certainly common, as the virus is endemic in most high-density cattle-raising regions. Since BCoVs are phylogenetically close to SARS-CoV-2, it is possible that cross-protection against COVID-19 occurs in people exposed to BCoV. MethodsThis article shows an in silico investigation of human cross-protection to SARS-CoV-2 due to BCoV exposure. We determined HLA recognition and human B lymphocyte reactivity to BCoV epitopes using bioinformatics resources. A retrospective geoepidemiological analysis of COVID-19 was then performed to verify if BCoV/SARS-CoV-2 cross-protection could have occurred in the field. Brazil was used as a model for the epidemiological analysis of the impact of livestock density - as a proxy for human exposure to BCoV - on the prevalence of COVID-19 in people. ResultsAs could be expected from their classification in the same Betacoronavirus genus, we show that several human B and T epitopes are shared between BCoV and SARS-CoV-2. This raised the possibility of cross-protection of people from exposure to the bovine coronavirus. Analysis of field data added partial support to the hypothesis of viral cross-immunity from human exposure to BCoV. There was a negative correlation between livestock geographical density and COVID-19. Whole-Brazil data showed areas in the country in which COVID-19 prevalence was disproportionally low (controlled by normalization by transport infrastructure). Areas with high cattle density had lower COVID-19 prevalence in these low-risk areas. ConclusionsThese data are hypothesis-raising indications that cross-protection is possibly being induced by human exposure to the Bovine Coronavirus.

10
Hypothetical Human Immune Genome Complex Gradient May Help To Explain The Congenital Zika Symdrome Catastrophe In Brazil: A New Theory

Oliveira, F.; Bresani-Salvi, C.; Morais, C.; Bigham, A.; Braga-Neto, U.; Maestre, G.; Vandebergh, J. L.; Marques, E.; Acioli-Santos, B.

2020-06-04 genetics 10.1101/2020.06.03.132878 medRxiv
Top 0.1%
18.9%
Show abstract

There are few data considering human genetics as an important risk factor for birth abnormalities related to ZIKV infection during pregnancy, even though sub-Saharan African populations are apparently more resistant to CZS as compared to populations in the Americas. We hypothesized that single nucleotide variants (SNVs), especially in innate immune genes, could make some populations more susceptible to Zika congenital complications than others. Differences in the SNV frequencies among continental populations provide great potential for Machine Learning techniques. We explored a key immune genomic gradient between individuals from Africa, Asia and Latin America, working with complex signatures, using 297 SNVs. We employed a two-step approach. In the first step, decision trees (DTs) were used to extract the most discriminating SNVs among populations. In the second step, machine learning algorithms were used to evaluate the quality of the SNV pool identified in step one for discriminating between individuals from sub-Saharan African and Latin-American populations. Our results suggest that 10 SNVs from 10 genes (CLEC4M, CD58, OAS2, CD80, VEPH1, CTLA4, CD274, CD209, PLAAT4, CREB3L1) were able to discriminate sub-Saharan Africans from Latin American populations using only immune genome data, with an accuracy close to 100%. Moreover, we found that these SNVs form a genome gradient across the three main continental populations. These SNVs are important elements of the innate immune system and in the response against viruses. Our data support the Human Immune Genome Complex Gradient hypothesis as a new theory that may help to explain the CZS catastrophe in Brazil.

11
Structural Homology and Electrostatic Potential Comparisons of Epitope Pair Candidates for Molecular Mimicry Triggering of Type 1 Diabetes Mellitus

Gardner, R.; Wilkins, J.; Mistry, S.; Gouripeddi, R.; Facelli, J.

2025-09-04 immunology 10.1101/2025.08.29.673145 medRxiv
Top 0.1%
18.8%
Show abstract

BackgroundMolecular mimicry, where foreign and self-peptides contain similar epitopes, can induce autoimmune responses. Identifying potential molecular mimics and studying their properties is key to understanding the onset of autoimmune diseases such as type 1 diabetes mellitus (T1DM). Previous work identified pairs of infectious epitopes (EINF) and T1DM epitopes (ET1D) that demonstrated sequence homology; however, structural homology was not considered. Correlating sequence homology with structural properties is important for translational investigation of potential molecular mimics. This work compares sequence homology with structural homology by calculating the structures and electrostatic potential surfaces of the epitope pairs identified in previous work from our laboratory. ResultsFor each pair of EINF and ET1D, the root mean square deviations (RMSD) were calculated between their predicted structures and their electrostatic potentials. Structures were predicted using the AlphaFold software program. Of the 53 epitope pairs considered here, only 10 did not exhibit any matching (i.e. less than 3 residues overlap). When considering all residues the RMSD ranges from 0.33 [A] to 11.66 [A] with an average of 2.68 [A]. Twenty-two pairs (42%) have RMSD of less than 1.5 [A] and 30 (58%) less than 3 [A]. ConclusionsMost of the EINF/ET1D pairs selected by sequence homology show similar structural and electrostatic distributions, indicating that the EINF may also bind to the same protein targets, i.e. the major histocompatibility complex molecules, for T1DM, leading to molecular mimicry onset of the disease. These findings suggest that searching for epitope pairs using sequence homology, a much less computationally demanding approach, leads to strong candidates for molecular mimicry that should be considered for further study. But structure homology, electrostatic potential calculations and full docking calculations may be necessary to advance the in-silico molecular mimicry predictions, which may be useful to select the most promising candidates for experimental studies.

12
Conserved T-cell epitopes predicted by bioinformatics in SARS-COV-2 variants

Lu, F.; Wang, S.; Wang, Y.; Yao, Y.; Wang, Y.; Liu, S.; Wang, Y.; Yu, Y.; Wang, L.

2021-08-13 immunology 10.1101/2021.08.12.456182 medRxiv
Top 0.1%
18.8%
Show abstract

BackgroundFinding conservative T cell epitopes in the proteome of numerous variants of SARS-COV-2 is required to develop T cell activating SARS-COV-2 capable of inducing T cell responses against SARS-COV-2 variants. MethodsA computational workflow was performed to find HLA restricted CD8+ and CD4+ T cell epitopes among conserved amino acid sequences across the proteome of 474727 SARS-CoV-2 strains. ResultsA batch of covserved regions in the amino acid sequences were found in the proteome of the SARS-COV-2 strains. 2852 and 847 peptides were predicted to have high binding affinity to distint HLA class I and class II molecules. Among them, 1456 and 484 peptides are antigenic. 392 and 111 of the antigenic peptides were found in the conseved amino acid sequences. Among the antigenic-conserved peptides, 6 CD8+ T cell epitopes and 7 CD4+ T cell epitopes were identifed. The T cell epitopes could be presented to T cells by high-affinity HLA molecules which are encoded by the HLA alleles with high population coverage. ConclusionsThe T cell epitopes are conservative, antigenic and HLA presentable, and could be constructed into SARS-COV-2 vaccines for inducing protective T cell immunity against SARS-COV-2 and their variants.

13
Immunoinformatics approach for a novel multi-epitope vaccine construct against spike protein of human coronaviruses

kumar, A.; Rathi, E.; Kini, S. G.

2021-05-02 immunology 10.1101/2021.05.02.442313 medRxiv
Top 0.1%
18.6%
Show abstract

Spike (S) proteins are an attractive target as it mediates the binding of the SARS-CoV-2 to the host through ACE-2 receptors. We hypothesize that the screening of S protein sequences of all the HCoVs would result in the identification of potential multi-epitope vaccine candidates capable of conferring immunity against various HCoVs. In the present study, several machine learning-based in-silico tools were employed to design a broad-spectrum multi-epitope vaccine candidate against S protein of human coronaviruses. To the best of our knowledge, it is one of the first study, where multiple B-cell epitopes and T-cell epitopes (CTL and HTL) were predicted from the S protein sequences of all seven known HCoVs and linked together with an adjuvant to construct a potential broad-spectrum vaccine candidate. Secondary and tertiary structures were predicted, validated and the refined 3D-model was docked with an immune receptor. The vaccine candidate was evaluated for antigenicity, allergenicity, solubility, and its ability to achieve high-level expression in bacterial hosts. Finally, the immune simulation was carried out to evaluate the immune response after three vaccine doses. The designed vaccine is antigenic (with or without the adjuvant), non-allergenic, binds well with TLR-3 receptor and might elicit a diverse and strong immune response.

14
Molecular epidemiology of the globally spreading genetic lineage IV of peste des petits ruminants virus

Courcelles, M.; Tounkara, K.; Mantip, S.; Niang, M.; Kounta Sidibe, C. A.; Sery, A.; Dakouo, M.; Luka, P. D.; Adedeji, A.; Shamaki, D.; Muhammad, M.; Ali, Y. H.; Saeed, I. K.; Awuni, J.; Odoom, T.; Tetteh, P. A.; Yingar, D. T.; Wade, A.; Dickmu, S.; Diddi, A.; Shawash, H.; Couacy-Hymann, E.; Mathurin, K. Y.; Ouled Ahmed Ben Ali, H.; Ben Hassen, S.; hadouchi, s.; Alm-ajali, A.; Settypalli, T. B. K.; Lamien, C. E.; Salami, H.; Rassoul, S.; Asnaoui, M.; Cetre-Sossah, C.; Guendouz, S.; Kwiatek, O.; Libeau, G.; Dundon, W. G.; Bataille, A.

2026-05-18 evolutionary biology 10.64898/2026.05.18.725933 medRxiv
Top 0.1%
18.6%
Show abstract

Peste des petits ruminants (PPR) is a highly contagious viral disease of small ruminants caused by the peste des petits ruminants virus (PPRV), which is classified into four distinct genetic lineages (I-IV). A critical concern in the recent epidemiological history of PPRV is the rapid and widespread expansion of lineage IV (LIV) across West Africa over the past decade. This dominance suggests a potential adaptive advantage of circulating LIV strains in the regions current epidemiological context. In this study, we obtain the genome sequence of 26 new PPRV samples, including historical (pre-2000) and many recent African LIV isolates, offering the first opportunity to investigate the evolutionary history of LIV in Africa and identify genetic events potentially associated with its recent spread. Phylogenomic analyses implemented on a dataset of 167 curated PPRV genome sequences reveal that the most ancestral LIV group comprises strains circulating in Sub-Saharan Africa (designated clade LIVssa), providing robust evidence for an African origin of lineage IV. Our results further indicate that PPRV strains linked to the recent West African expansion of LIV belong to a specific LIVssa subgroup, termed NigB. We identified multiple signatures of selection pressure within the LIVssa sublineage, particularly in the NigB cluster. Several amino acid substitutions unique to LIVssa or NigB were detected, some of which may impact protein function and warrant prioritised investigation. Additional genomic data are required to confirm the association between the NigB group and the ongoing spread of LIV in West Africa. The evolutionary adaptations observed in LIVssa - potentially enhancing transmission efficiency, host range or pathogenicity - could undermine current disease control strategies in regions where PPR poses significant threats to food security and local economies. Author SummaryPeste des petits ruminants virus (PPRV) infects sheep and goats across Africa, Middle East, Asia and Europe, causing disease with major impact on global economy and food security. One genetic lineage of PPRV, called lineage IV (LIV), is at the origin of most recent expansion of the distribution of the disease, including replacement of other lineages in areas of African where PPRV is historically present. Here, we generated genome sequences from PPRV LIV isolates from different dates and places to study the evolution of this genetic lineage and explore whether its recent spread can be associated with the appearance of new mutations in the virus genome. Our results provide evidence that the PPRV LIV originated in Sub-Saharan Africa and identify mutations present only virus isolates currently spready in new regions of Africa. Further research should investigate the impact of these mutations on protein functions and capacity of transmission of PPRV.

15
Whole Genome Sequencing-based Characterization of Human Genome Variation and Mutation Burden in Botswana

Thami, P. K.; Choga, W. T.; Mulisa, D. D.; Dandara, C.; Shevchenko, A. K.; Leteane, M. M.; Novitsky, V.; O'Brien, S. J.; Essex, M.; Gaseitsiwe, S.; Chimusa, E. R.

2020-12-15 genomics 10.1101/2020.12.15.422821 medRxiv
Top 0.1%
18.3%
Show abstract

The study of human genome variations can contribute towards understanding population diversity and the genetic aetiology of health-related traits. We sought to characterise human genomic variations of Botswana in order to assess diversity and elucidate mutation burden in the population using whole genome sequencing. Whole genome sequences of 390 unrelated individuals from Botswana were available for computational analysis. The sequences were mapped to the human reference genome GRCh38. Population joint variant calling was performed using Genome Analysis Tool Kit (GATK) and BCFTools. Variant characterisation was achieved by annotating the variants with a suite of databases in ANNOVAR and snpEFF. The genomic architecture of Botswana was delineated through principal component analysis, structure analysis and FST. We identified a total of 27.7 million unique variants. Variant prioritisation revealed 24 damaging variants with the most damaging variants being ACTRT2 rs3795263, HOXD12 rs200302685, ABCB5 rs111647033, ATP8B4 rs77004004 and ABCC12 rs113496237. We observed admixture of the Khoe-San, Niger-Congo and European ancestries in the population of Botswana, however population substructure was not observed. This exploration of whole genome sequences presents a comprehensive characterisation of human genomic variations in the population of Botswana and their potential in contributing to a deeper understanding of population diversity and health in Africa and the African diaspora.

16
Mutational signatures in countries affected by SARS-CoV-2: Implications in host-pathogen interactome

Rahman, S. A.; Singh, J.; Singh, H.; Hasnain, S. E.

2020-09-17 bioinformatics Community evaluation 10.1101/2020.09.17.301614 medRxiv
Top 0.1%
18.3%
Show abstract

We are in the midst of the third severe coronavirus outbreak caused by SARS-CoV-2 with unprecedented health and socio-economic consequences due to the COVID-19. Globally, the major thrust of scientific efforts has shifted to the design of potent vaccine and anti-viral candidates. Earlier genome analyses have shown global dominance of some mutations purportedly indicative of similar infectivity and transmissibility of SARS-CoV-2 worldwide. Using high-quality large dataset of 25k whole-genome sequences, we show emergence of new cluster of mutations as result of geographic evolution of SARS-CoV-2 in local population ([≥]10%) of different nations. Using statistical analysis, we observe that these mutations have either significantly co-occurred in globally dominant strains or have shown mutual exclusivity in other cases. These mutations potentially modulate structural stability of proteins, some of which forms part of SARS-CoV-2-human interactome. The high confidence druggable host proteins are also up-regulated during SARS-CoV-2 infection. Mutations occurring in potential hot-spot regions within likely T-cell and B-cell epitopes or in proteins as part of host-viral interactome, could hamper vaccine or drug efficacy in local population. Overall, our study provides comprehensive view of emerging geo-clonal mutations which would aid researchers to understand and develop effective countermeasures in the current crisis. SignificanceOur comparative analysis of globally dominant mutations and region-specific mutations in 25k SARS-CoV-2 genomes elucidates its geo-clonal evolution. We observe locally dominant mutations (co-occurring or mutually exclusive) in nations with contrasting COVID-19 mortalities per million of population) besides globally dominant ones namely, P314L (ORF1b) and D164G (S) type. We also see exclusive dominant mutations such as in Brazil (I33T in ORF6 and I292T in N protein), England (G251V in ORF3a), India (T2016K and L3606F in ORF1a) and in Spain (L84S in ORF8). The emergence of these local mutations in ORFs within SARS-CoV-2 genome could have interventional implications and also points towards their potential in modulating infectivity of SARS-CoV-2 in regional population.

17
SARS-CoV-2 genome analysis of strains in Pakistan reveals GH, S and L clade strains at the start of the pandemic

Ghanchi, N. K.; Masood, K. I.; Nasir, A.; Khan, W.; Abidi, S. H.; Shahid, S.; Mahmood, S. F.; Kanji, A. R.; Razzak, S. A.; Ansar, Z.; Islam, N.; Dharejo, M. B.; Hasan, Z.; Hasan, R.

2020-08-04 genomics 10.1101/2020.08.04.234153 medRxiv
Top 0.1%
18.2%
Show abstract

ObjectivesPakistan has a high infectious disease burden with about 265,000 reported cases of COVID-19. We investigated the genomic diversity of SARS-CoV-2 strains and present the first data on viruses circulating in the country. MethodsWe performed whole-genome sequencing and data analysis of SARS-CoV-2 eleven strains isolated in March and May. ResultsStrains from travelers clustered with those from China, Saudi Arabia, India, USA and Australia. Five of eight SARS-CoV-2 strains were GH clade with Spike glycoprotein D614G, Ns3 gene Q57H, and RNA dependent RNA polymerase (RdRp) P4715L mutations. Two were S (ORF8 L84S and N S202N) and three were L clade and one was an I clade strain. One GH and one L strain each displayed Orf1ab L3606F indicating further evolutionary transitions. ConclusionsThis data reveals SARS-CoV-2 strains of L, G, S and I have been circulating in Pakistan from March, at the start of the pandemic. It indicates viral diversity regarding infection in this populous region. Continuing molecular genomic surveillance of SARS-CoV-2 in the context of disease severity will be important to understand virus transmission patterns and host related determinants of COVID-19 in Pakistan.

18
Higher frequency of interstate over international transmission chains of SARS-CoV-2 virus at the Rio Grande do Sul - Brazil state borders

Dezordi, F. Z.; Silva Junior, J. V. J.; Ruoso, T. F.; Batista, A. G.; Fonseca, P. M.; Salvato, R. S.; Gregianini, T. S.; Lopes, T. R. R.; Flores, E. F.; Weiblen, R.; Brites, P. C.; Silva, M. d. M.; da Rocha, J. B. T.; Barbosa, G. d. L.; Machado, L. C.; da Silva, A. F.; Paiva, M. H. S.; Bezerra, M. F.; Campos, T. d. L.; Gräf, T.; Sganzerla, D. A.; Loreto, E. L. d. S.; Wallau, G. d. L.

2024-05-21 infectious diseases 10.1101/2024.05.21.24307668 medRxiv
Top 0.1%
17.7%
Show abstract

Brazils COVID-19 response has faced challenges due to the continuous emergence of variants of concern (VOCs), emphasizing the need for ongoing genomic surveillance and retrospective analyses of past epidemic waves. Rio Grande do Sul (RS), Brazils southernmost state, has crucial international borders and trades with Argentina and Uruguay, along with significant domestic connections. The source and sink of transmission with both national and international hubs raises questions about the RS role in the transmission of the virus, which has not been fully explored. Nasopharyngeal samples from various municipalities in RS were collected between June 2020 and July 2022. SARS-CoV-2 whole genome amplification and sequencing were performed using high-throughput Illumina sequencing. Bioinformatics analysis encompassed the development of scripts and tools to take into account epidemiological information to reduce sequencing disparities bias among the regions/countries, genome assembly, and large-scale phylogenetic reconstruction. Here, we sequenced 1,480 SARS-CoV-2 genomes from RS, covering all major regions. Sequences predominantly represented Gamma (April-June 2021) and Omicron (January-July 2022) variants. Phylogenetic analysis revealed a regional pattern for transmission dynamics, particularly with Southeast Brazil for Gamma, and a range of inter-regional connections for Delta and Omicron within the country. On the other hand, international and cross-border transmission with Argentina and Uruguay was rather limited. We evaluated the three VOCs circulation over two years in RS using a new subsampling strategy based on the number of cases in each state during the circulation of each VOC. In summary, the retrospective analysis of genomic surveillance data demonstrated that virus transmission was less intense between country borders than within the country. These findings suggest that while non-pharmacological interventions were effective to mitigate transmission across international land borders in RS, they were unsuficient to contain transmission at the domestic level.

19
Immunoinformatic Designing and Evaluation of a Broad-Spectrum Multiepitope Vaccine Against MDR Acinetobacter baumannii, Klebsiella pneumoniae and Pseudomonas aeruginosa

Bhardwaj, N.; Sah, S. N.; Billa, A.; Gupta, V.; Capalash, N.; Sharma, P.

2025-06-01 immunology 10.1101/2025.05.27.656513 medRxiv
Top 0.1%
16.0%
Show abstract

Acinetobacter baumannii, Klebsiella pneumoniae and Pseudomonas aeruginosa are among the multidrug-resistant (MDR) Gram-negative pathogens that pose a growing threat, necessitating novel preventive measures in addition to traditional antibiotics. By using advanced immunoinformatics methods, highly conserved and immunogenic epitopes for B-cells and T-cells were selected from major virulence-associated proteins (LptE, YiaD, MrkD, PhoE, OprF, and Zot), which exhibited high antigenicity (VaxiJen scores 0.74-2.75) and 98.87% global population coverage. The construct vaccine comprises 50S ribosomal protein L7/L12 adjuvant and PADRE sequence for immunogenicity enhancement, with structural validation indicating stability (96.3% residues in the total allowed Ramachandran regions). High-affinity interactions with TLR2/TLR4 (binding energies: -1009.6 to -1079.6 kcal/mol) were found through molecular docking, and immune simulations suggested strong humoral (IgM/IgG) and cellular (IFN-{gamma}/IL-12) responses. Importantly, MEP vaccines can overcome major drawbacks of traditional vaccines by (1) offering cross-strain protection via conserved epitopes, (2) lowering the need for antibiotics through infection prevention, and (3) providing affordable options for healthcare systems affected by MDR infections. These findings demonstrate the MEP constructs potential as a preventative measure against nosocomial infections, which may have implications for combating the global AMR epidemic. Further experimental validation is needed to verify its efficacy.

20
Mutation Landscape of SARS COV2 in Africa

Nassir, A. A.; Musanabaganwa, C.; Mwikarago, I.

2020-12-21 genomics 10.1101/2020.12.20.423630 medRxiv
Top 0.1%
15.2%
Show abstract

COVID-19 disease has had a relatively less severe impact in Africa. To understand the role of SARS CoV2 mutations on COVID-19 disease in Africa, we analysed 282 complete nucleotide sequences from African isolates deposited in the NCBI Virus Database. Sequences were aligned against the prototype Wuhan sequence (GenBank accession: NC_045512.2) in BWA v. 0.7.17. SAM and BAM files were created, sorted and indexed in SAMtools v. 1.10 and marked for duplicates using Picard v. 2.23.4. Variants were called with mpileup in BCFtools v. 1.11. Phylograms were created using Mr. Bayes v 3.2.6. A total of 2,349 single nucleotide polymorphism (SNP) profiles across 294 sites were identified. Clades associated with severe disease in the United States, France, Italy, and Brazil had low frequencies in Africa (L84S=2.5%, L3606F=1.4%, L3606F/V378I/=0.35, G251V=2%). Sub Saharan Africa (SSA) accounted for only 3% of P323L and 4% of Q57H mutations in Africa. Comparatively low infections in SSA were attributed to the low frequency of the D614G clade in earlier samples (25% vs 67% global). Higher disease burden occurred in countries with higher D614G frequencies (Egypt=98%, Morocco=90%, Tunisia=52%, South Africa) with D614G as the first confirmed case. V367F, D364Y, V483A and G476S mutations associated with efficient ACE2 receptor binding and severe disease were not observed in Africa. 95% of all RdRp mutations were deaminations leading to CpG depletion and possible attenuation of virulence. More genomic and experimental studies are needed to increase our understanding of the temporal evolution of the virus in Africa, clarify our findings, and reveal hot spots that may undermine successful therapeutic and vaccine interventions.