Concatenation of segmented viral genomes for reassortment analysis
Ivanova, A. O.; Volchkov, P. Y.; Deviatkin, A. A.
Show abstract
Most reassortment identification methods are based on searching for phylogenetic discrepancies between phylogenetic trees for different segments. Other methods use pairwise genetic distances or compare the position of individual genome components in a tree relative to a reference component of the viral genome. However, such approaches are labour-intensive and hardly scalable. Recent advances in the availability of viral sequencing technologies have led to the sequencing of large numbers of pathogen genomes, making manual processing of this large data difficult. At the same time, recombination analysis methods can process almost any number of sequences simultaneously. Such approaches are not suitable for the simultaneous analysis of multiple segments and are therefore not used to search for reassortment events. However, in the case of sequential concatenation of all segments, the methods of recombination analysis can be used to detect traces of reassortment events. The code is available at https://github.com/melibrun/Concatenation-of-segmented-viral-genomes-for-reassortment-analysis. The service is implemented as a web application https://melibrun.shinyapps.io/viralsegmentconcatenator1/. It concatenates segmented viral genomes for reassortment analysis. The tool accepts files in GenBank format as input and generates a set of sequences in fasta format that are sequentially concatenated sequences of viral segments named in accordance with the "strain" field of the GenBank record annotation. In order to use recombination search algorithms in the study of reassortment events, we have developed a method (Virus Segment Concatenator, VSC) to automatically concatenate the sequences of all segments of a virus into a single sequence. The applicability of VSC for automated searches for reassortment events was demonstrated using CCHFV, an H5N5 subtype of influenza virus.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An advanced sequence clustering and designation workflow reveals the enzootic maintenance of a dominant West Nile virus subclade in Germany 96%
- Unique protein features of SARS-CoV-2 relative to other Sarbecoviruses 95%
- Begomovirus species demarcation based on genome-sequence identity often yields non-monophyletic species: A case study of sweet potato-infecting begomoviruses 95%
Similar papers in this journal
- Machine learning using intrinsic genomic signatures for rapid classification of novel pathogens: COVID-19 case study 95%
- Exploring G and C-quadruplex structures as potential targets against the severe acute respiratory syndrome coronavirus 2 94%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 94%
Similar papers in this journal
- In-depth Bioinformatic Analyses of Human SARS-CoV-2, SARS-CoV, MERS-CoV, and Other Nidovirales Suggest Important Roles of Noncanonical Nucleic Acid Structures in Their Lifecycles 95%
- Positive selection of ORF3a and ORF8 genes drives the evolution of SARS-CoV-2 during the 2020 COVID-19 pandemic 95%
- Comparative Genomics and Environmental Distribution of Large dsDNA viruses in the family Asfarviridae 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.