Genomic Similarity of Nucleotides in SARS CoronaVirus using K-Means Unsupervised Learning Algorithm
Singh, J.
Show abstract
The drastic increase in the number of coronaviruses discovered and coronavirus genomes being sequenced have given us a great opportunity to perform genomics and bioinformatics analysis on this family of viruses. Coronaviruses possess the largest genomes (26.4 to 31.7 kb) among all known RNA viruses, with G + C contents varying from 32% to 43%. Phylogenetically, three genera, Alphacoronavirus, Betacoronavirus and Gammacoronavirus, with Betacoronavirus consisting of subgroups A, B, C were known to exist but now a new genus D also exists,namely the Deltacoronavirus. In such a situation, it becomes highly important for efficient classification of all virus data so that it helps us in suitable planning,containment and treatment. The objective of this paper is to classify SARS corona-virus nucleotide sequences based on parameters such as sequence length,percentage similarity between the sequence information,open and closed gaps in the sequence due to multiple mutations and many others.By doing this,we will be able to predict accurately the similarity of SARS CoV-2 virus with respect to other corona-viruses like the Wuhan corona-virus,the bat corona-virus and the pneumonia virus and would help us better understand about the taxonomy of the corona-virus family. SUMMARYIn addition to the guidelines provided in the abstract above,the following points summarizes the article below: O_LIThe article discusses an application of Machine Learning in the field of virology. C_LIO_LIIt aims to classify the SARS CoV2 virus as per the already known sequences of the bat-coronavirus, the Wuhan Sea Food Market pneumonia virus and the Wuhan coronavirus. C_LIO_LITo solve and predict the similarity of the SARS CoV2 coronavirus w.r.t other viruses discussed above,K-Means Unsupervised Learning Algorithm has been chosen. C_LIO_LIThe data-set used is MN997409.1-4NY0T82X016-Alignment-HitTable.csv found on www.kaggle.com.(Complete link shared in the references section).[17] C_LIO_LIThe results have been validated by using a simple data-correlation technique namely Spearmans Rank Correlation Coeffecient. C_LIO_LII have also discussed my future work using Deep Neural Nets that can help predict new virus sequences and effectively find similarity if any with already discovered viruses. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial intelligence tool for the study of COVID-19 microdroplet spread across the human diameter and airborne space 96%
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 96%
- GenomeBits insight into omicron and delta variants of coronavirus pathogen 95%
Similar papers in this journal
- Predicting the Epidemic Curve of the Coronavirus (SARS-CoV-2) Disease (COVID-19) Using Artificial Intelligence 95%
- Viral miRNAs Confer Survival in Host Cells by Targeting Apoptosis Related Host Genes 94%
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 94%
Similar papers in this journal
- Prediction of high-risk liver cancer patients from their mutation profile: Benchmarking of mutation calling techniques 93%
- MINTyper: An outbreak-detection method for accurate and rapid SNP typing of clonal clusters with noisy long reads 91%
- Convolutional-LSTM Approach for Temporal Catch Hotspots (CATCH): An AI-Driven Model for Spatiotemporal Forecasting of Fisheries Catch Probability Densities 91%
Similar papers in this journal
- Detection of spreader nodes and ranking of interacting edges in Human-SARS-CoV protein interaction network 95%
- An Issue of Concern: Unique Truncated ORF8 Protein Variants of SARS-CoV-2 95%
- A machine learning approach for identification of gastrointestinal predictors for the risk of COVID-19 related hospitalization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.