PRIZM: Combining Low-N Data and Zero-shot Models to Design Enhanced Protein Variants
Harding-Larsen, D.; Lax, B. M.; Garcia, M. E.; Mendonca, C.; Mejia-Otalvaro, F.; Welner, D. H.; Mazurenko, S.
Show abstract
Machine learning has repeatedly shown the ability to accelerate protein engineering, but many approaches demand large amounts of robust, high-quality training data as well as substantial computational expertise. While large pre-trained models can function as zero-shot proxies for predicting variant effects, selecting the best model for a given protein property is often non-trivial. Here, we introduce Protein Ranking using Informed Zero-shot Modelling (PRIZM), a two-phase workflow that first uses a high-quality low-N dataset to identify the most suitable pre-trained zero-shot model for a target protein property and then applies that model to rank and prioritize an in silico variant library for experimental testing. Across diverse benchmark datasets spanning multiple protein properties, PRIZM reliably separated low-from high-performing models using datasets of [~]20 labelled variants. We further demonstrate PRIZM in enzyme engineering case studies targeting sucrose synthase thermostability and glycosyltransferase activity, where PRIZM-guided selection identified improved variants, including gains of [~]3{degrees}C in apparent melting temperature and [~]20% higher relative activity. PRIZM provides an accessible, data-efficient route to leverage foundation models for protein design while requiring minimal experimental data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Predicts New Anti-CRISPR Proteins 94%
- Beyond DNA Binding: single C2H2 zinc fingers with adjacent β-strands mediate dimerization in Drosophila transcription factors 94%
- Structural and functional characterization of DdrC, a novel DNA damage-induced nucleoid associated protein involved in DNA compaction 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.