Back

High-sensitivity whole-genome recovery of single viral species in environmental samples

Chen, L.; Chen, A.; Zhang, X. D.; Saenz, M. T.; Han, H.-S.; Xiao, Y.; Xiao, G.; Pipas, J. M.; Weitz, D.

2023-11-16 bioengineering
10.1101/2023.11.13.566948 bioRxiv
Show abstract

Characterizing unknown viruses is essential for understanding viral ecology and preparing against viral outbreaks. Recovering complete genome sequences from environmental samples remains computationally challenging using metagenomics, especially for low-abundance species with uneven coverage. This work presents a method for reliably recovering complete viral genomes from complex environmental samples. Individual genomes are encapsulated into droplets and amplified using multiple displacement amplification. A novel gene detection assay, which employs an RNA-based probe and an exonuclease, selectively identifies droplets containing the target viral genome. Labeled droplets are sorted using a microfluidic sorter, and genomes are extracted for sequencing. Validation experiments using a sewage sample spiked with two known viruses demonstrate the methods efficacy. We achieve 100% recovery of the spiked-in SV40 (Simian virus 40, 5243bp) genome sequence with uniform coverage distribution, and approximately 99.4% for the larger HAd5 genome (Human Adenovirus 5, 35938bp). Notably, genome recovery is achieved with as few as one sorted droplet, which enables the recovery of any desired genomes in complex environmental samples, regardless of their abundance. This method enables targeted characterizations of rare viral species and whole-genome amplification of single genomes for accessing the mutational profile in single virus genomes, contributing to an improved understanding of viral ecology.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.