Data augmentation enables label-specific generation of homologous protein sequences
Rosset, L.; Weigt, M.; Zamponi, F.
Show abstract
Accurately annotating and controlling protein function from sequence data remains a major challenge, particularly within homologous families where annotated sequences are scarce and structural variation is minimal. We present a two-stage approach for semi-supervised functional annotation and conditional sequence generation in protein families using representation learning. First, we demonstrate that protein language models, pretrained on large and diverse sequence datasets and possibly finetuned via contrastive learning, provide embeddings that robustly capture fine-grained functional specificities, even with limited labeled data. Second, we use the inferred annotations to train a generative probabilistic model, an annotation-aware Restricted Boltzmann Machine, capable of producing synthetic sequences with prescribed functional labels. Across several protein families, we show that this approach achieves highly accurate annotation quality and supports the generation of functionally coherent sequences. Our findings underscore the power of combining self-supervised learning with light supervision to overcome data scarcity in protein function prediction and design.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Correlations from structure and phylogeny combine constructively in the inference of protein partners from sequences 97%
- Phylogenetic correlations can suffice to infer protein partners from sequences 97%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.