Back

parSEQ: PROBE AND RESCUE SEQUENCING FOR ADVANCED VARIANT RETRIEVAL FROM DNA POOL

Houmani, M.; Peterkin, F.; Antoun, G.; Fischer, L.; Hammi, A.

2023-12-12 molecular biology
10.1101/2023.12.12.571337 bioRxiv
Show abstract

Modern protein engineering is powered by sequence-function data-sets. We have developed parSEQ, a platform that maximizes the capture of these protein sequence-function data-sets through a sequence-first-screen-later approach. parSEQ relies on Next-Generation Sequencing (NGS) to reverse the conventional screen-first-sequence-later workflow. This allows for the high-throughput retrieval of variant sequences from DNA pools, ensuring that every screened variants functional data is paired with its sequence data. This report details parSEQs methodology and its integration into various protein engineering workflows. Through several case studies, we illustrate parSEQs broad applicability. These case studies describe the use of parSEQ for high-throughput variant DNA-template retrieval, expression-ready bacterial clone isolation, sourcing DNA for denovo designed proteins, and preparing targeted mutational libraries. Our findings suggest that parSEQs approach to capturing sequence-function data-sets can advance protein engineering efforts, particularly in the age of artificial intelligence and machine learning.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.