A graph representation of gapped patterns in phage sequences for graph convolutional network
WANG, R.; NG, Y. K.; ZHANG, X.; WANG, J.; Li, S.
Show abstract
Genome sequencing technologies reveal a huge amount of genomic sequences. Neural network-based methods can be prime candidates for retrieving insights from these sequences because of their applicability to large and diverse datasets.However, the highly variable lengths of nucleic acid sequences severely impair the presentation of sequences as input to the neural network. Genetic variations further complicate tasks that involve sequence comparison or alignment. Here, we propose a graph representation of nucleic acid sequences called gapped pattern graphs. These graphs can be transformed through a Graph Convolutional Network to form lower-dimensional embeddings for downstream tasks. On the basis of the gapped pattern graphs, we implemented a neural network model and demonstrated its performance in studying phage sequences. We compared our model with equivalent models based on other forms of input in performing four tasks related to nucleic acid sequences--phage and ICE discrimination, phage integration site prediction, lifestyle prediction, and host prediction. Other state-of-the-art tools were also compared, where available. Our method consistently outperformed all the other methods in various metrics on all four tasks. In addition, our model was able to identify distinct gapped pattern signatures from the sequences.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 95%
- PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings 95%
- JIND: Joint Integration and Discrimination for Automated Single-Cell Annotation 95%
Similar papers in this journal
- Improved metagenomic analysis with Kraken 2 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.