Viral Sequence Identification in Metagenomes using Natural Language Processing Techniques
Abdelkareem, A. O.; Khalil, M. I.; Elbehery, A. H.; Abbas, H. M.
Show abstract
Viral reads identification is one of the important steps in metagenomic data analysis. It shows up the diversity of the microbial communities and the functional characteristics of microorganisms. There are various tools that can identify viral reads in mixed metagenomic data using similarity and statistical tools. However, the lack of available genome diversity is a serious limitation to the existing techniques. In this work, we applied natural language processing approaches for document classification in analyzing metagenomic sequences. Text featurization is presented by treating DNA similar to natural language. These techniques reveal the importance of using the text feature extraction pipeline in sequence identification by transforming DNA base pairs into a set of characters with a term frequency and inverse document frequency techniques. Various machine learning classification algorithms are applied to viral identification tasks such as logistic regression and multi-layer perceptron. Moreover, we compared classical machine learning algorithms with VirFinder and VirNet, our deep attention model for viral reads identification on generated fragments of viruses and bacteria for benchmarking viral reads identification tools. Then, as a verification of our tool, It was applied to a simulated microbiome and virome data for tool verification and real metagenomic data of Roche 454 and Illumina for a case study.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 94%
- MACI: A machine learning-based approach to identify drug classes of antibiotic resistance genes from metagenomic data 94%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.