PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
Arora, R. K.; Chen, L. T.; Du, M.; Marks, D.; Church, G.
Show abstract
General-purpose language models are being increasingly utilized in protein-design workflows, yet their ability to evaluate variant effects remains unclear. To answer this question, we introduce PG-LLM, a benchmark built on ProteinGym to evaluate general-purpose language models on 217 protein-variant prioritization tasks. Each task follows the same format: a language model receives a wild-type protein sequence, an assay description, and is tasked with ranking 50 mutant sequences by fitness without access to multiple-sequence alignments or protein structures. We evaluate thirteen language models and rescore 95 published protein predictors on the same candidate sets with the same evaluation metric. Claude Opus 5 leads the primary leaderboard with a Spearman correlation of{rho} = 0.406, narrowly ahead of GPT-5.6 Sol at 0.402. However, GPT-5.6 Sol scores higher than Opus 5 when the two models are compared only on assays scored by both. Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at{rho} = 0.411, but remains below the leading predictor VenusREM at{rho} = 0.523. Variant-ranking performance improves with test-time compute across GPT, Claude, and Gemini models, but the gains taper before closing the gap to specialist protein predictors. Unlike sequence-only predictors, which perform better on proteins with deeper evolutionary alignments, LLM accuracy changes little across alignment-depth. To address contamination risk, we additionally source 59 DMS assays from 19 studies whose scores first became public after January 2026. On this set, we evaluate three LLMs and seven biomolecular baselines, observing performance trends similar to those on the main PG-LLM benchmark. PG-LLM shows that tool-free language models capture substantial protein-variant signal and already outperform many established sequence-based predictors. These results establish the emerging capability of language models as biomolecular reasoners while defining the remaining headroom for their reliable use in variant-prioritization workflows.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 94%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 94%
- ProtNote: a multimodal method for protein-function annotation 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 95%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 94%
- ProteinMCP: An Agentic AI Framework for Autonomous Protein Engineering 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.