Back

Developing a Standard Definition for Sequences of Concern

Alexanian, T.; Beal, J.; Bartling, C.; Berlips, J.; Carr, P. A.; Clore, A.; Cozzarini, H.; Diggans, J.; El Moubayed, Y.; Esvelt, K.; Flyangolts, K.; Foner, L.; Fullerton, P. A.; Gemler, B. T.; Jagla, C. A.; Lababidi, R.; Mitchell, T.; Murphy, S. T.; Parker, M. T.; Roehner, N.; Rusch, A.; Talley, K.; Timmerman, T.; Wheeler, N. E.

2026-03-18 bioinformatics
10.64898/2026.03.14.711820 bioRxiv
Show abstract

Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential "sequences of concern," there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a test set of 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequence of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the test set as a whole. Together, these results provide a concrete "sequence of concern" definition that can be used as a foundation for development of biosecurity screening standards and policy.

Published in Frontiers in Bioengineering and Biotechnology (predicted rank #7) · training set

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.