Back

Quantifying uncertainty in Protein Representations Across Models and Task

Prabakaran, R.; Bromberg, Y.

2025-05-06 bioinformatics
10.1101/2025.04.30.651545 bioRxiv
Show abstract

Embeddings, derived by language models, are widely used as numeric proxies for human language sentences and structured data. In the realm of biomolecules, embeddings serve as efficient sequence and/or structure representations, enabling similarity searches, structure and function prediction, and estimation of biophysical and biological properties. However, relying on embeddings without assessing the models confidence in its ability to accurately represent molecular properties is a critical flaw--akin to using a scalpel in surgery without verifying its sharpness. In this study, we propose a means to evaluate the ability of protein language models to represent proteins, assessing their capacity to encode biologically relevant information. Our findings reveal that low-quality embeddings often fail to capture meaningful biology, displaying vector properties indistinguishable from those of randomly generated sequences. A key contributor to this performance issue is the models failure to learn the underlying biology from unevenly distributed sequence spaces in the training data. Our novel, model-agnostic scoring framework is, to the best of our knowledge, the first to quantify protein sequence embedding reliability. We believe that our robust approach to screening embeddings prior to making biological inferences, stands to significantly enhance the reliability of downstream applications.

Published in Nature Methods (predicted rank #6) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.