Back

Breaking the Synthesis Barrier for AI-Designed DNA Libraries

Sussex, S.; Borevkovic, E.; Lohmann, F.; Chen, N.; Lüthi, E.; Reddy, S. T.; Krause, A.

2026-07-07 bioengineering
10.64898/2026.07.07.736931 bioRxiv
Show abstract

Designing DNA libraries is a key challenge from drug design to protein engineering and synthetic biology. Modern generative models offer opportunities to navigate the design space and propose specific sequences predicted to be effective in-silico. Designing deterministic libraries of specific sequences is however limited by the cost of DNA synthesis -- the synthesis barrier. In contrast, high-throughput multiplexed screening can measure the function of billions of biological sequences in parallel. Harnessing this technology requires the design of randomized libraries with specific design constraints to achieve low synthesis costs. In practice, such stochastic libraries are often chosen heuristically, sacrificing control for scale. Is there a way to bridge AI-based in-silico sequence design with high-throughput experimentation? In this work, we introduce Policy Gradients for Library Design (PGLD). PGLD uses a synthesis-aware parametrization of stochastic DNA libraries and optimizes them against a specified objective function. This allows for designing massive, controlled libraries without being limited by synthesis costs. We show how PGLD enables lab-in-the-loop design of multi-round high-throughput experiments, and large-scale in-vitro DNA sampling from generative models. Finally, we use PGLD to design a library of ~10^6 unique sequences which is synthesized at a cost of ~700 USD to explore the mutation space of a broadly neutralizing influenza antibody.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
38.3%
2
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 5%
7.6%
3
Nature Communications
5641 papers in training set
Top 22%
7.6%
50% of probability mass above
4
PLOS Computational Biology
1863 papers in training set
Top 6%
6.5%
5
Nature Biotechnology
172 papers in training set
Top 1%
3.1%
6
eLife
5828 papers in training set
Top 36%
3.1%
7
Nature Methods
385 papers in training set
Top 3%
2.6%
8
Nature Machine Intelligence
70 papers in training set
Top 1%
2.4%
9
ACS Synthetic Biology
287 papers in training set
Top 1%
2.1%
10
Science Advances
1243 papers in training set
Top 18%
1.8%
11
Genome Research
468 papers in training set
Top 4%
1.7%
12
Neuron
337 papers in training set
Top 4%
1.7%
13
Nature Computational Science
55 papers in training set
Top 0.7%
1.7%
14
Science
477 papers in training set
Top 5%
1.7%
15
iScience
1154 papers in training set
Top 19%
1.6%
16
Journal of The Royal Society Interface
235 papers in training set
Top 3%
1.4%
17
Scientific Reports
3612 papers in training set
Top 68%
1.1%
18
Nature Biomedical Engineering
47 papers in training set
Top 1%
1.0%
19
Genome Biology
637 papers in training set
Top 8%
0.9%
20
Cell
431 papers in training set
Top 11%
0.8%
21
Nature
645 papers in training set
Top 11%
0.8%
22
Nucleic Acids Research
1281 papers in training set
Top 14%
0.8%
23
Trends in Biotechnology
12 papers in training set
Top 0.4%
0.6%
24
Nature Neuroscience
252 papers in training set
Top 6%
0.6%
25
Cell Reports Methods
165 papers in training set
Top 5%
0.6%