Protein Dimension DB: A Unified Protein Repository for Representation Learning and Functional Analysis
Alves Sobrinho, P. d. A.; Sakamoto, T.; Figuerola, W. B.
Show abstract
Inspired by the success of large language models in areas like natural language processing, researchers have applied similar architectures, notably the Transformer, to protein sequences. Thanks to these developments, Protein Language Models (PLMs) have become important resources for diverse tasks such as predicting protein family, function, solubility, cellular location, molecular interactions and remote homology. However, the size of the best performing PLMs (which can be up to 15B parameters) requires substantial computational power. Protein Dimension DB addresses this critical bottleneck by providing a centralized, version-controlled resource of precomputed protein embeddings, experimentally validated molecular function annotations, and taxonomic encodings. The database integrates embeddings from seven state-of-the-art PLMs, including ProtT5, ESM2, and Ankh variants for all Swiss-Prot/ UniProt proteins. These models were compared by benchmarking molecular function prediction. Tests revealed that hybrid embeddings (e.g., Ankh Base + ProtT5) outperformed single-model approaches with minimal dimensionality increases. Taxonomic encodings further boosted performance by 2.9% AUPRC, demonstrating lineage-aware learning. By providing embeddings in Parquet format -- a columnar storage optimized for machine learning workflows -- the resource eliminates GPUdependent preprocessing and reduces storage requirements. This enables immediate use in resource-constrained environments while maintaining backward compatibility through versioned releases. All datasets are freely accessible via Github and HuggingFace, with unified metadata enabling applications from functional annotation to evolutionary studies. Protein Dimension DB bridges the gap between cutting-edge PLMs and practical biological research, offering researchers standardized inputs for reproducible, multi-modal protein analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 96%
- InterLabelGO+: Unraveling label correlations in protein function prediction 95%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 95%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 95%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 94%
- NERVE 2.0: boosting the New Enhanced Reverse Vaccinology Environment via artificial intelligence and a user-friendly web interface 94%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 93%
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.