Negligible peptidome diversity of SARS-CoV-2 and its higher taxonomic ranks
Chong, L. C.; Khan, A. M.
Show abstract
The unprecedented increase in SARS-CoV-2 sequence data limits the application of alignment-dependent approaches to study viral diversity. Herein, we applied our recently published UNIQmin, an alignment-free tool to study the protein sequence diversity of SARS-CoV-2 (sub-species) and its higher taxonomic lineage ranks (species, genus, and family). Only less than 0.5% of the reported SARS-CoV-2 protein sequences are required to represent the inherent viral peptidome diversity, which only increases to a mere [~]2% at the family rank. This is expected to remain relatively the same even with further increases in the sequence data. The findings have important implications in the design of vaccines, drugs, and diagnostics, whereby the number of sequences required for consideration of such studies is drastically reduced, short-circuiting the discovery process, while still providing for a systematic evaluation and coverage of the pathogen diversity.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Analyses of phylogenetics, natural selection, and protein structure of clade 2.3.4.4b H5N1 Influenza A reveal that recent viral lineages have evolved promiscuity in host range and improved replication in mammals in North America 93%
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 92%
- Unraveling the molecular basis of host cell receptor usage in SARS-CoV-2 and other human pathogenic β-CoVs 91%
Similar papers in this journal
- Identification and quantification of SARS-CoV-2 leader subgenomic mRNA gene junctions in nasopharyngeal samples shows phasic transcription in animal models of COVID-19 and dysregulation at later time points that can also be identified in humans 93%
- IPEV: Identification of Prokaryotic and Eukaryotic Virus-derived sequences in virome using deep learning 93%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.