Decoding the physicochemical basis of taxonomy preferences in protein design models
Dillon, L. B.; Crook, O. M.; Maiwald, A.
Show abstract
Protein design models have transformed protein engineering by enabling computational exploration of sequence spaces far exceeding experimental capacity. However, their outputs are shaped by both the protein distributions represented in their training corpora and the information available during scoring, so the same model score may reflect backbone-compatible biophysics, taxonomic structure in sequence databases, or other learned regularities rather than protein fitness alone. Here we quantify systematic preferences across 14 protein design models that differ in data modality, training-corpus composition, and scoring context, for comparison grouped as backbone-conditioned, structure plus native-sequence context, or sequence-only. Backbone-conditioned models retain little unexplained taxonomic variance after controlling for protein family and measurable biophysical properties, with residual species variance below 3.3%. In contrast, sequence-only models retain substantial residual taxonomic dependence of 15-20%, indicating that likelihood remains strongly entangled with organism-level sequence statistics. These differences across model classes produce distinct preference landscapes. Backbone-conditioned models organise scores around compactness, packing, and charge, while sequence-only models preserve stronger within-family taxonomic effects. Redesign experiments show that these preferences propagate into generation, shifting templates toward characteristic biophysical profiles rather than uniformly sampling backbone-compatible sequence space. Continued training of ProteinMPNN on ecologically selected extremophile secretomes redirects designed surface chemistry along an acid-base axis while largely preserving structural compatibility and global taxonomic structure. These results show that systematic preferences are not a single failure mode, but separable components arising from scoring context, training-corpus composition, and learned biophysical constraints. Together, they provide a framework for disentangling the sources of model preference and linking them to both scoring behaviour and generated sequence properties.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ConforFold Recovers Alternative Protein Conformations Beyond MSA Subsampling 94%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 94%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.