Deacon: fast sequence filtering and contaminant depletion
Constantinides, B.; Lees, J.; Crook, D. W.
Show abstract
MotivationRealising the value of large DNA sequence collections demands efficient search and extraction of sequences of interest. Search queries may vary in size from short gene sequences to multiple whole genomes that are too large to fit in computer memory. In microbial genomics, a routine search application involving both large queries and large collections is the removal of contaminating host genome sequences from microbial (meta)genomes. Where the host is human, sensitive classification and excision of host sequences is usually necessary to protect host genetic information. Precise classification is also critical in order to retain microbial sequences and permit accurate microbial genomic analysis. While human pangenomes have been shown to increase sensitivity of human sequence classification, existing bioinformatic host depletion approaches have either limited precision when used with metagenomes or large computing resource requirements. ResultsWe present Deacon, an efficient and versatile sequence filter for raw sequence files and streams. We demonstrate its leading accuracy for the task of host depletion, using less computing resource than existing approaches. By querying a human pangenome index for minimizers contained in each input sequence, Deacon is able to accurately classify and discard diverse human sequences from long reads at over 250Mbp/s with a commodity laptop. We present validation of classification sensitivity, specificity and speed with simulated short and long reads for diverse catalogues of human, bacterial and viral genomes alongside existing methods. Beyond host depletion, Deacon is well suited to common sequence search and filtering applications, particularly those involving large queries. Capable of indexing a human genome in under 30s, Deacon is equipped to rapidly compose custom minimizer indexes using set operations, facilitating efficient search and filtering of massive sequence datasets using gigabase queries. Availability and implementationDeacon is implemented as an MIT-licensed command line tool written in Rust and packaged with Bioconda. Code is available from https://github.com/bede/deacon.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
- iGenomics: Comprehensive DNA Sequence Analysis on your Smartphone 95%
- Pangenome databases provide superior host removal and mycobacteria classification from clinical metagenomic data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.