DETIRE: A Hybrid Deep Learning Model for identifying Viral Sequences from Metagenomes
Miao, Y.; Liu, F.; Hou, T.; Liu, Q.; Dong, T.; Liu, Y.
Show abstract
A metagenome contains all DNA sequences from an environmental sample, including viruses, bacteria, fungi, actinomycetes and so on. Since viruses are of huge abundance and have caused vast mortality and morbidity to human society in history as a kind of major pathogens, detecting viruses from metagenomes plays a crucial role in analysing the viral component of samples and is the very first step for clinical diagnosis. However, detecting viral fragments directly from the metagenomes is still a tough issue because of the existence of huge number of short sequences. In this paper, a hybrid Deep lEarning model for idenTifying vIral sequences fRom mEtagenomes (DETIRE), is proposed to solve the problem. Firstly, the graph-based nucleotide sequence embedding strategy is utilized to enrich the expression of DNA sequences by training an embedding matrix. Then the spatial and sequential features are extracted by trained CNN and BiLSTM networks respectively to improve the feature expression of short sequences. Finally, the two set of features are weighted combined for the final decision. Trained by 220,000 sequences of 500bp subsampled from the Virus and Host RefSeq genomes, DETIRE identifies more short viral sequences (<1,000bp) than three latest methods, DeepVirFinder, PPR-Meta and CHEER. DETIRE is freely available at https://github.com/crazyinter/DETIRE.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LSTM-PHV: Prediction of human-virus protein-protein interactions by LSTM with word2vec 97%
- VIGA: an one-stop tool for eukaryotic Virus Identification and Genome Assembly from next-generation-sequencing data 96%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 96%
Similar papers in this journal
Similar papers in this journal
- Prediction of virus-host association using protein language models and multiple instance learning 97%
- Deep6mA: a deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species 96%
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.