Efficient algorithms for designing maximally sized orthogonal DNA sequence libraries
Gowri, G.; Sheng, K.; Yin, P.
Show abstract
Orthogonal sequence library design is an essential task in bioengineering. Typical design approaches scale quadratically in the size of the candidate sequence space. As such, exhaustive searches of sequence space to maximize library size are computationally intractable with existing methods. Here, we present SeqWalk, a time and memory efficient method for designing maximally-sized orthogonal sequence libraries using the sequence symmetry minimization heuristic. SeqWalk encodes sequence design constraints in a de Bruijn graph representation of sequence space, enabling the application of efficient graph traversal techniques to the problem of orthogonal DNA sequence design. We demonstrate the scalability of SeqWalk by designing a provably maximal set of > 106 orthogonal 25nt sequences in less than 20 seconds on a single standard CPU core. We additionally derive fundamental bounds on orthogonal sequence library size under a variety of design constraints.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Mutation rate variations in the human genome are encoded in DNA shape 94%
- LazySampling and LinearSampling: Fast Stochastic Sampling of RNA Secondary Structure with Applications to SARS-CoV-2 94%
- RoboCOP: Jointly computing chromatin occupancy profiles for numerous factors from chromatin accessibility data 93%
Similar papers in this journal
- Single-sequence protein-RNA complex structure prediction by geometric attention-enabled pairing of biological language models 92%
- RepairSig: Deconvolution of DNA damage and repaircontributions to the mutational landscape of cancer 91%
- Inferring protein sequence-function relationships with large-scale positive-unlabeled learning 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.