Developing a Standard Definition for Sequences of Concern
Alexanian, T.; Beal, J.; Bartling, C.; Berlips, J.; Carr, P. A.; Clore, A.; Cozzarini, H.; Diggans, J.; El Moubayed, Y.; Esvelt, K.; Flyangolts, K.; Foner, L.; Fullerton, P. A.; Gemler, B. T.; Jagla, C. A.; Lababidi, R.; Mitchell, T.; Murphy, S. T.; Parker, M. T.; Roehner, N.; Rusch, A.; Talley, K.; Timmerman, T.; Wheeler, N. E.
Show abstract
Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential "sequences of concern," there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a test set of 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequence of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the test set as a whole. Together, these results provide a concrete "sequence of concern" definition that can be used as a foundation for development of biosecurity screening standards and policy.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The META tool optimizes metagenomic analyses across sequencing platforms and classifiers. 92%
- Early Detection of Emerging SARS-CoV-2 Variants of Interest for Experimental Evaluation 92%
- Using short-read 16S rRNA sequencing of multiple variable regions to generate high-quality results to a species level. 90%
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 93%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 92%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 92%
Similar papers in this journal
- Rapid and accurate SNP genotyping of clonal bacterial pathogens with BioHansel 92%
- Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high throughput identification and lineage assignment in Pakistan 91%
- Development of a nextflow bioinformatics pipeline for the detection of SARS-CoV-2 co-infection cases from genomic surveillance in the Philippines 91%
Similar papers in this journal
- Understanding Ecological Systems Using Knowledge Graphs: An Application to Highly Pathogenic Avian Influenza 92%
- Craft: A Machine Learning Approach to Dengue Subtyping 92%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.