Back

Decoding the physicochemical basis of taxonomy preferences in protein design models

Dillon, L. B.; Crook, O. M.; Maiwald, A.

2026-07-23 biochemistry
10.1101/2025.10.21.683350 bioRxiv
Show abstract

Protein design models have transformed protein engineering by enabling computational exploration of sequence spaces far exceeding experimental capacity. However, their outputs are shaped by both the protein distributions represented in their training corpora and the information available during scoring, so the same model score may reflect backbone-compatible biophysics, taxonomic structure in sequence databases, or other learned regularities rather than protein fitness alone. Here we quantify systematic preferences across 14 protein design models that differ in data modality, training-corpus composition, and scoring context, for comparison grouped as backbone-conditioned, structure plus native-sequence context, or sequence-only. Backbone-conditioned models retain little unexplained taxonomic variance after controlling for protein family and measurable biophysical properties, with residual species variance below 3.3%. In contrast, sequence-only models retain substantial residual taxonomic dependence of 15-20%, indicating that likelihood remains strongly entangled with organism-level sequence statistics. These differences across model classes produce distinct preference landscapes. Backbone-conditioned models organise scores around compactness, packing, and charge, while sequence-only models preserve stronger within-family taxonomic effects. Redesign experiments show that these preferences propagate into generation, shifting templates toward characteristic biophysical profiles rather than uniformly sampling backbone-compatible sequence space. Continued training of ProteinMPNN on ecologically selected extremophile secretomes redirects designed surface chemistry along an acid-base axis while largely preserving structural compatibility and global taxonomic structure. These results show that systematic preferences are not a single failure mode, but separable components arising from scoring context, training-corpus composition, and learned biophysical constraints. Together, they provide a framework for disentangling the sources of model preference and linking them to both scoring behaviour and generated sequence properties.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.