VirRep: accurate identification of viral genomes from human gut metagenomic data via a hybrid language representation learning framework
Dong, Y.; Chen, W.; Zhao, X.
Show abstract
Accurate identification of viral genomes from metagenomic data provides a broad avenue for studying viruses in the human gut. Here, we introduce VirRep, a novel virus identification method based on a hybrid language representation learning framework. VirRep employs a context-aware encoder and a composition-focused encoder to incorporate the learned knowledge and known biological insights to better describe the source of a DNA sequence. We benchmarked VirRep on multiple human gut virome datasets under different conditions and demonstrated significant superiority than state-of-the-art methods and even combinations of them. A comprehensive validation has also been conducted on real human gut metagenomes to show the great utility of VirRep in identifying high-quality viral genomes that are missed by other methods.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 96%
- stGCL: A versatile cross-modality fusion method based on multi-modal graph contrastive learning for spatial transcriptomics 95%
- High-precision cell-type mapping and annotation of single-cell spatial transcriptomics with STAMapper 95%
Similar papers in this journal
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- vRhyme enables binning of viral genomes from metagenomes 95%
- Probabilistic tensor decomposition extracts better latent embeddings from single-cell multiomic data 95%
Similar papers in this journal
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 96%
- Deep autoregressive generative models capture the intrinsics embedded in T-cell receptor repertoires 96%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.