Scalable and cost-efficient custom gene library assembly from oligopools
Freschlin, C.; Yang, K.; Romero, P. A.
Show abstract
Advances in metagenomics, deep learning, and generative protein design have enabled broad in silico exploration of sequence space, but experimental characterization is still constrained by the cost and scalability of DNA synthesis. Here, we present OMEGA (Oligo-based Multiplexed Efficient Gene Assembly), a low-cost, accessible method for assembling hundreds to thousands of full-length genes in parallel using standard laboratory techniques. OMEGA computationally fragments target genes into short, high-fidelity Golden Gate-compatible oligonucleotides that can be ordered as a pooled library and assembled across multiplexed subpools. We systematically optimized the number of fragments per gene and orthogonal ligation sites per reaction and determine that OMEGA can assemble up to 2.6 kb constructs using as many as 70 Golden Gate sites. To validate the approach, we assembled and functionally screened a library of 810 natural and synthetic GFP variants, recovering 94-97% of target sequences with high uniformity. OMEGA enables precision library construction at scale, with per-gene costs as low as $1.50, and offers a broadly applicable solution for bridging computational protein design with high-throughput experimental validation. We have developed OMEGA as an open-source software package and an easy-to-use Colab notebook available at https://github.com/RomeroLab/omega.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Genome-wide functional screens enable the prediction of high activity CRISPR-Cas9 and -Cas12a guides in Yarrowia lipolytica 96%
- A split ribozyme that links detection of a native RNA to orthogonal protein outputs 95%
- Biochemical-free enrichment or depletion of RNA classes in real-time during direct RNA sequencing with RISER 95%
Similar papers in this journal
- Probing molecular specificity with deep sequencing and biophysically interpretable machine learning 94%
- Nanopore adaptive sequencing for mixed samples, whole exome capture and targeted panels. 94%
- Highly accurate barcode and UMI error correction using dual nucleotide dimer blocks allows direct single-cell nanopore transcriptome sequencing 94%
Similar papers in this journal
- ConSeqUMI, an error-free nanopore sequencing pipeline to identify and extract individual nucleic acid molecules from heterogeneous samples 96%
- Arrayed in vivo barcoding for multiplexed sequence verification of plasmid DNA and demultiplexing of pooled libraries 95%
- DropSynth 2.0: high-fidelity multiplexed gene synthesis in emulsions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.