Quantifying uncertainty in Protein Representations Across Models and Task
Prabakaran, R.; Bromberg, Y.
Show abstract
Embeddings, derived by language models, are widely used as numeric proxies for human language sentences and structured data. In the realm of biomolecules, embeddings serve as efficient sequence and/or structure representations, enabling similarity searches, structure and function prediction, and estimation of biophysical and biological properties. However, relying on embeddings without assessing the models confidence in its ability to accurately represent molecular properties is a critical flaw--akin to using a scalpel in surgery without verifying its sharpness. In this study, we propose a means to evaluate the ability of protein language models to represent proteins, assessing their capacity to encode biologically relevant information. Our findings reveal that low-quality embeddings often fail to capture meaningful biology, displaying vector properties indistinguishable from those of randomly generated sequences. A key contributor to this performance issue is the models failure to learn the underlying biology from unevenly distributed sequence spaces in the training data. Our novel, model-agnostic scoring framework is, to the best of our knowledge, the first to quantify protein sequence embedding reliability. We believe that our robust approach to screening embeddings prior to making biological inferences, stands to significantly enhance the reliability of downstream applications.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 95%
- TemBERTure: Advancing protein thermostability prediction with Deep Learning and attention mechanisms 94%
- Mining hidden knowledge: Embedding models of cause-effect relationships curated from the biomedical literature 93%
Similar papers in this journal
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 95%
- Improving deep models of protein-coding potential with a Fourier-transform architecture and machine translation task 93%
- FilterDCA: interpretable supervised contact prediction using inter-domain coevolution 93%
Similar papers in this journal
- Disobind: a sequence-based, partner-dependent contact map and interface residue predictor for intrinsically disordered regions 95%
- Inferring protein sequence-function relationships with large-scale positive-unlabeled learning 95%
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.