Back

Accelerating protein design by scaling experimental characterization

Qian, J.; Milles, L. F.; Wicky, B. I. M.; Motmaen, A.; Li, X.; Kibler, R. D.; Stewart, L.; Baker, D.

2025-08-06 biochemistry
10.1101/2025.08.05.668824 bioRxiv
Show abstract

Recent advances in de novo protein design have greatly outpaced standard protein biochemistry workflows, and experimental testing has been a bottleneck in the validation of new designs and methodologies. Here, we describe experimental and computational workflows to address the issues of scale, speed and reproducibility of common in vitro protein testing methods, enabling at least an order of magnitude increase in throughput while reducing wetlab time. Semi-Automated Protein Production (SAPP) is a rapid, modular, scalable and cost-effective protocol, enabling up to milligram-scale protein production, and standardized characterization including yield, dispersity, and oligomeric state of hundreds of designs per day, at the cost-equivalent of a few DNA oligos per construct. End-to-end protocol execution takes 48 hours, with about 6 hours spent benchside using mostly standard laboratory equipment. This protocol has become the standard at our institute, providing critical experimental validation for dozens of projects spanning tens of thousands of designs. We showcase the power of the platform by using it to rapidly characterize de novo designed inhibitors of respiratory syncytial virus. Since at least 80% of SAPPs total cost comes from synthetic DNA, we also developed a scalable demultiplexing protocol (DMX) to leverage oligo pools as input DNA, providing a further 5-fold reduction in costs, enabling >1000 designs to be purified and characterized in arrayed, clonal format at a cost of $5 per construct. By reframing standard molecular biology practices and orchestrating wetlab workflows with partial automation instead of complex end-to-end robotics, these protocols should be widely adoptable, accelerating protein design.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.