Optimizing Consensus Generation Algorithms for Highly Variable Amino Acid Sequence Clusters
Mohabati, R.; Rezaei, R.; Mohajel, N.; Ranjbar, M. M.; Azadmanesh, K.; Roohvand, F.
Show abstract
Producing a functional consensus sequence is a preliminary bioinformatics task, which is a necessity for many research purposes. However, the existence of hypervariable regions in the input multiple sequence alignment files causes complications in generating a useful consensus sequence. The current methods for consensus generation, Threshold, and majority algorithms, have several problems, which exclude them as applicable algorithms for such highly variable sequence clusters. Hence, we designed a novel alternative algorithm for the same purpose. The algorithm was explained both using a mathematical formula and a practical implementation in Python programming language. A sequence set from HCV genotype 1b E2 protein has been utilized as a practical example for evaluating the algorithms performance. A few in silico tests have been performed on the output sequence and the results have been compared to results from other algorithms. Epitope-mapping analysis indicates the functionality of this algorithm, by preserving the hotspot residues in the consensus sequence, and the antigenicity index shows significant antigenicity of the consensus sequence. Moreover, phylogenetic analysis shows no significant change in the placement of the new consensus sequence on the phylogenetic tree compared to other algorithms. This approach will have several implications in designing a new vaccine for highly variable viruses such as HIV-1, Influenza, and Hepatitis C Viruses (HCV).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- iBRAB: in silico based-designed Broad-spectrum Fab against H1N1 Influenza A Virus 96%
- Designing of a next generation multiepitope based vaccine (MEV) against SARS-COV-2: Immunoinformatics and in silico approaches 96%
- In silico comparative genomics of SARS-CoV-2 to determine the source and diversity of the pathogen in Bangladesh 96%
Similar papers in this journal
- Design of multi epitope-based peptide vaccine against E protein of human COVID-19: An immunoinformatics approach 96%
- Proteochemometric method for pIC50 prediction of Flaviviridae 95%
- Expression of nitric oxide synthase and nitric oxide levels in peripheral blood cells and oxidized low-density lipoprotein levels in saliva as early markers of severe dengue 93%
Similar papers in this journal
- Epitope-based peptide vaccine against glycoprotein G of Nipah henipavirus using immunoinformatics approaches 97%
- Design of Epitope Based Peptide Vaccine Against Pseudomonas Aeruginosa Fructose Bisphosphate Aldolase Protein using Immunoinformatics 97%
- Attenuated Subcomponent Vaccine Design Targeting the SARS-CoV-2 Nucleocapsid Phosphoprotein RNA Binding Domain: In silico analysis 96%
Similar papers in this journal
- Identification of novel mutations in RNA-dependent RNA polymerases of SARS-CoV-2 and their implications on its protein structure 97%
- An Issue of Concern: Unique Truncated ORF8 Protein Variants of SARS-CoV-2 97%
- A machine learning approach for identification of gastrointestinal predictors for the risk of COVID-19 related hospitalization 95%
Similar papers in this journal
- Epitope-Based Peptide Vaccine against Bombali Ebolavirus Viral Protein 40: An Immunoinformatics Combined with Molecular Docking Studies 97%
- A Multiple Peptides Vaccine against nCOVID-19 Designed from the Nucleocapsid phosphoprotein (N) and Spike Glycoprotein (S) via the Immunoinformatics Approach 96%
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.