Development of a fast feature extraction method for SARS-CoV-2 spike sequences using amino acid physicochemical properties
Oka, H.; Noba, K.; Sasahara, J.; Hashimoto, T.; Yoshimoto, S.; Niioka, H.; Miyake, J.; Hori, K.
Show abstract
COVID-19 continues to spread today, leading to an accumulation of SARS-CoV-2 virus mutations in databases, and large amounts of genomic datasets are currently available. However, due to these large datasets, utilizing this amount of sequence data without random sampling is challenging. Major difficulties for downstream analyses include the increase in the dimension size along with the conversion of sequences into numerical values when using conventional amino acid representation methods, such as one-hot encoding and k-mer-based approaches that directly reflect sequences. Moreover, these sequences are deficient in physicochemical characteristics, such as structural information and hydrophilicity; hence, they fail to accurately represent the inherent function of the given sequences. In this study, we utilized the physicochemical properties of amino acids to develop a rapid and efficient approach for extracting feature parameters that are suitable for downstream processes of machine learning, such as clustering. A fixed-length feature vector representation of a spike sequence with reduced dimensionality was obtained by converting amino acid residues into physicochemical parameters. Next, t-distributed stochastic neighbor embedding (t- SNE), a method for dimensionality reduction and visualization of high-dimensional data, was performed, followed by density-based spatial clustering of applications with noise (DBSCAN). The results show that by using the physicochemical properties of amino acids rather than conventional methods that directly represent sequences into numerical values, SARS-CoV-2 spike sequences can be clustered with sufficient accuracy and a shorter runtime. Interestingly, the clusters obtained by using amino acid properties include subclusters that are distinct from those produced utilizing the method for the direct representation of amino acid sequences. A more detailed analysis indicated that the contributing parameters of this novel cluster identified exclusively when utilizing the physicochemical properties of amino acids significantly differ from one another. This suggests that representing amino acid sequences by physicochemical properties might enable the identification of clusters with enhanced sensitivity compared to conventional methods. Author summaryOne of the major causes of the global threat of SARS-CoV-2 is the rapid emergence of its variants. While analyzing these variants is crucial for understanding the mechanism of outbreaks, the expansion of database size is becoming a barrier for effective analysis. In this study, we provide an approach that allows researchers without vast computational resources to comprehensively analyze the variants of SARS-CoV-2 spike by representing the sequences using the physicochemical properties of amino acids. The result of clusters derived using this method demonstrates not only an accuracy comparable to the conventional approaches of directly converting sequences into numerical values but also indicates the potential for more detailed clustering outcomes. The results suggest that our approach is valuable for the rapid identification of characteristic residues in new variants of SARS-CoV-2 and other viruses that may arise in the future.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 96%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 95%
Similar papers in this journal
- Bioinformatics analysis and collection of protein post-translational modification sites in human viruses 96%
- DBpred: A deep learning method for the prediction of DNA interacting residues in protein sequences 96%
- An in silico approach to identification, categorization and prediction of nucleic acid binding proteins 96%
Similar papers in this journal
Similar papers in this journal
- SARS-CoV-2 NSP14 governs mutational instability and assists in making new SARS-CoV-2 variants 96%
- MACI: A machine learning-based approach to identify drug classes of antibiotic resistance genes from metagenomic data 96%
- A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence 95%
Similar papers in this journal
- AutoVEM2: a flexible automated tool to analyze candidate key mutations and epidemic trends for virus 95%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 95%
- Real-time monitoring epidemic trends and key mutations in SARS-CoV-2 evolution by an automated tool 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.