The Protein Language Visualizer: Sequence Similarity Networks for the Era of Language Models
Espinoza Herrera, J.; Manriquez Garcia, M. F.; Medina Bermejo, S.; Lopez Jasso, A.; Shi, K.; Mead, D.; Veskimägi, S. M.; O'Connor, M.; Siordia, A.; Roethler, N.; Jinich, A.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe era of modern AI-driven representations of proteins is here, and moving fast, yet tools for their intuitive visualization and exploration lag behind. Sequence Similarity Networks (SSNs) have long filled this role for alignment-based methods, providing simple but widely adopted platforms for grouping proteins by homology. Building on this foundation, we present the Protein Language Visualizer (PLVis), a modular framework that applies existing pre-trained protein language model (pLM) embeddings, dimensionality reduction, and clustering to generate interactive maps of protein relationships. The central contribution is the PLVis Repository, an online resource where thousands of reference proteomes can be compared and annotated through an accessible, interactive interface, much like SSNs became impactful not for their technical novelty but for their broad usability. We first validate that well-separated clusters in PLVis reliably capture homology information, while emphasizing caution when interpreting central "fuzzy" regions. We then illustrate the value of PLVis through case studies spanning individual protein families to full proteome comparisons across Mycobacterium and Plasmodium species. By combining methodological clarity with broad accessibility, the PLVis Repository provides a low-barrier platform for exploring proteomes through the lens of language models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 95%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 94%
Similar papers in this journal
- Efficient indexing of peptides for database search using Tide 94%
- METATRYP v 2.0: Metaproteomic Least Common Ancestor Analysis for Taxonomic Inference Using Specialized Sequence Assemblies - Standalone Software and Web Servers for Marine Microorganisms and Coronaviruses 94%
- Fast and memory efficient searching of large-scale mass spectrometry data using Tide 94%
Similar papers in this journal
- FlatProt: 2D visualization eases protein structure comparison 95%
- NERVE 2.0: boosting the New Enhanced Reverse Vaccinology Environment via artificial intelligence and a user-friendly web interface 94%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
Similar papers in this journal
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 94%
- Limits and potential of combined folding and docking using PconsDock. 94%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.