Back

Generating Hybrid Proteins with the MSA-Transformer

Tule, S.; Davis, S.; Koludarov, I.; Mora, A.; Boden, M.

2026-03-10 bioinformatics
10.1101/2025.11.20.689447 bioRxiv
Show abstract

Protein superfamilies display extensive sequence and functional divergence, providing a rich landscape for engineering functionally enhanced variants. We present a stochastic, iterative framework that leverages the MSA-Transformer to generate intermediate sequences between a homologous pair of user-specified "source" and "target" proteins in sequence space. Sequence sites to "mask" is either selected by embedding-based dissimilarity or by row-attention information, while beam search concurrently explores multiple mutational pathways. Pre-trained sparse autoencoders combined with sequence and structural analyses are used to trace the inheritance and exchange of features across the "mutational pathway", revealing "hybrid" sequences that integrate properties of both source and target proteins. Applied across diverse protein families, the framework produces sequences that occupy biologically meaningful regions of sequence space and achieve higher consistency and plausibility scores than random baselines according to in silico metrics. In the B1/B2 metallo-{beta}-lactamase family, hybrids largely retain their core fold recombining structural and active-site motifs from both subclasses, demonstrating the models capacity to preserve catalytic features while exploring novel structural permutations. Code availabilityThe implementation is available via GitHub at https://github.com/santule/protmixy.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.