Studying Pathogens Degrades BLAST-based Pathogen Identification
Beal, J.; Clore, A.; Manthey, J.
Show abstract
As synthetic biology becomes increasingly capable and accessible, it is likewise increasingly critical to be able to make accurate biosecurity determinations regarding the pathogenicity or toxicity of particular nucleic acid or amino acid sequences. At present, this is typically done using the BLAST algorithm to determine the best match with sequences in the NCBI databases. Neither BLAST nor the NCBI databases, however, are actually designed for biosafety determination. Critically, taxonomic errors or ambiguities in the NCBI databases can also cause errors in BLAST-based taxonomic categorization. With heavily studied taxa and frequently used biotechnology tools, even low frequency taxonomic categorization issues can lead to high rates of errors in biosecurity decision-making. Here we focus on the implications for false positives, finding that NCBI BLAST will now incorrectly categorize a number of commonly used biotechnology tool sequences as the pathogens or toxins with which they have been used. Paradoxically, this implies that problems are expected to be most acute for the pathogens and toxins of highest interest and the most widely used biotechnology tools. We thus conclude that biosecurity tools should shift away from BLAST against NCBI and towards new methods that are specifically tailored for biosafety purposes.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 94%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 93%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 93%
Similar papers in this journal
- Beyond Blast: Enabling Microbiologists to Better Extract Literature, Taxonomic Distributions and Gene Neighborhood Information for Protein Families 93%
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 92%
- Bakta: Rapid & standardized annotation of bacterial genomes via alignment-free sequence identification 92%
Similar papers in this journal
- Hypothesizing mechanistic links between microbes and disease using knowledge graphs 91%
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 91%
- K-PAM: A unified platform to distinguish Klebsiella species K- and O-antigen types, model antigen structures and identify hypervirulent strains 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.