Exploring protein sequence similarity with Protein Language UMAPs (PLUMAPs)
Jinich, A.; Nazia, S. Z.; Rhee, K. Y.
Show abstract
Visualizing relationships and similarities between proteins can reveal insightful biology. Current approaches to visualize and analyze proteins based on sequence homology, such as sequence similarity networks (SSNs), create representations of BLAST-based pairwise comparisons. These approaches could benefit from incorporating recent protein language models, which generate high-dimensional vector representations of protein sequences from self-supervised learning on hundreds of millions of proteins. Inspired by SSNs, we developed an interactive tool - Protein Language UMAPs (PLUMAPs) - to visualize protein similarity with protein language models, dimensionality reduction, and topic modeling. As a case study, we compare our tool to Sequence Similarity Network (SSN) using the proteomes of two related bacterial species, Mycobacterium tuberculosis and Mycobacterium smegmatis. Both SSNs and PLUMAPs generate protein clusters corresponding to protein families and highlight enrichment or depletion across species. However, only in PLUMAPs does the layout distance between proteins and protein clusters meaningfully reflect similarity. Thus in PLUMAPs, related protein families are displayed as nearby clusters, and larger-scale structures correlate with cellular localization. Finally, we adapt techniques from topic modeling to automatically annotate protein clusters, making them more easily interpretable and potentially insightful. We envision that as large protein language models permeate bioinformatics and interactive sequence analysis tools, PLUMAPs will become a useful visualization resource across a wide variety of biological disciplines. Anticipating this, we provide a prototype for an online, open source version of PLUMAPs.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 95%
- scFeatures: Multi-view representations of single-cell and spatial data for disease outcome prediction 93%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 93%
Similar papers in this journal
- EPIFANY - A method for efficient high-confidence protein inference 94%
- XlinkCyNET: a Cytoscape application for visualization of protein interaction networks based on cross-linking mass-spectrometry identifications 94%
- ProteoPlotter: an executable proteomics visualization tool compatible with Perseus 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.