Leveraging ancestral sequence reconstruction for protein representation learning
Matthews, D. M.; Spence, M. A.; Mater, A. C.; Nichols, J.; Pulsford, S. B.; Sandhu, M.; Kaczmarski, J. A. B.; Miton, C. M.; Tokuriki, N.; Jackson, C. J.
Show abstract
Protein language models (PLMs) convert amino acid sequences into the numerical representations required to train machine learning (ML) models. Many PLMs are large (>600 M parameters) and trained on a broad span of protein sequence space. However, these models have limitations in terms of predictive accuracy and computational cost. Here, we use multiplexed Ancestral Sequence Reconstruction (mASR) to generate small but focused functional protein sequence datasets for PLM training. Compared to large PLMs, this local ancestral sequence embedding (LASE) produces representations 10-fold faster and with higher predictive accuracy. We show that due to the evolutionary nature of the ASR data, LASE produces smoother fitness landscapes in which protein variants that are closer in fitness value become numerically closer in representation space. This work contributes to the implementation of ML-based protein design in real-world settings, where data is sparse and computational resources are limited.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accuracy and data efficiency in deep learning models of protein expression 97%
- Generalizable and scalable protein stability prediction with rewired protein generative models 97%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 97%
Similar papers in this journal
Similar papers in this journal
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 95%
- NEST: Spatially-mapped cell-cell communication patterns using a deep learning-based attention mechanism 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.