Decoding protein language models: insights from embedding space analysis
Rissom, P. F.; Yanez Sarmiento, P.; Safer, J.; Coley, C. W.; Renard, B. Y.; Heyne, H. O.; Iqbal, S.
Show abstract
Foundation models, which encode patterns in large, high-dimensional data as embeddings, show promise in many machine learning related applications in molecular biology. Embeddings learned by the models provide informative features for downstream prediction tasks, however, the information captured by the model is often not interpretable. One approach to understanding the captured information is through the analysis of their learned embeddings, which in molecular biology so far has mainly focused on visualizing individual embedding spaces. This study introduces a quantitative framework for cross-space comparison, enabling intuitive exploration and comparison of embedding spaces in molecular biology. The framework emphasizes analyzing the distribution of known biological information within embedding space neighborhoods and provides insights into relationships between multiple embedding spaces. Comparison techniques include global pairwise distance measurements as well as local nearest neighbor analyses. By applying our framework to embeddings from protein language models, we demonstrate how embedding space analysis can serve as a valuable pre-filtering step for task-specific supervised machine learning applications and for the recognition of differential patterns in data encoded within and across different embedding spaces. To support a wide usability, we provide a Python library that implements all analysis methods, available at https://github.com/broadinstitute/EmmaEmb.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
- Benchmarking Uncertainty Quantification for Protein Engineering 94%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
Similar papers in this journal
- DeepSS2GO: protein function prediction from secondary structure 94%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 93%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.