Multi-Modal Large Language Model Enables Protein Function Prediction
Huo, M.; Guo, H.; Cheng, X.; Singh, D.; Rahmani, H.; Li, S.; Gerlof, P.; Ideker, T.; Grotjahn, D. A.; Villa, E.; Song, L.; Xie, P.
Show abstract
Predicting the functions of proteins can greatly accelerate biological discovery and applications, where deep learning methods have recently shown great potential. However, these methods predominantly predict protein functions as discrete categories, which fails to capture the nuanced and complex nature of protein functions. Furthermore, existing methods require the development of separate models for each prediction task, a process that can be both resource-heavy and time-consuming. Here, we present ProteinChat, a versatile, multi-modal large language model that takes a proteins amino acid sequence as input and generates comprehensive narratives describing its function. ProteinChat is trained using over 1,500,000 (protein, prompt, answer) triplets curated from the Swiss-Prot dataset, covering diverse functions. This novel model can universally predict a wide range of protein functions, all within a single, unified framework. Furthermore, ProteinChat supports interactive dialogues with human users, allowing for iterative refinement of predictions and deeper exploration of protein functions. Our experimental results, evaluated through both human expert assessment and automated metrics, demonstrate that ProteinChat outperforms general-purpose LLMs like GPT-4, one of the flagship LLMs, by over ten-fold. In addition, ProteinChat exceeds or matches the performance of task-specific prediction models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 94%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 94%
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 97%
- Cracking the black box of deep sequence-based protein-protein interaction prediction 95%
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.