Back

Protein Language Modeling beyond static folds reveals sequence-encoded flexibility

Lueth, F. H.; Mihaila, V.; Mirdita, M.; Steinegger, M.; Rost, B.; Heinzinger, M.

2026-01-22 bioinformatics
10.64898/2026.01.21.700698 bioRxiv
Show abstract

MotivationProteins function through motion. Yet, most discoveries still commence with static representations of protein structures. Here, we investigated the feasibility of leveraging protein dynamics to improve homology detection. ResultsWe introduce ProtProfileMD, a sequence-to-3D-probability model that predicts, from an amino acid sequence, a profile of discrete structural representations capturing protein dynamics. We applied supervised parameter-efficient finetuning of the ProstT5 protein Language Model (pLM) to predict per-residue distributions over Foldseeks 3Di alphabet derived from motions observed in molecular dynamics. This original result reveals that the 3Di tokens, despite being coarse-grained descriptors of 3D structure, still offer sufficient resolution to capture aspects of conformational changes. This is evidenced by a correlation between fluctuations in the 3D protein structure over the course of a molecular dynamics trajectory and the entropy of 3Di states. Based on this insight, we introduce a proof-of-concept for making remote homology detection of proteins more sensitive by leveraging a proteins distinctive dynamic fingerprint captured by our model. Our method recovers flexibility signals with a fidelity that is biologically relevant, improving search and complementing protein structure predictions, for example, by flagging flexible, disordered, or other functionally relevant regions. Availability and ImplementationProtProfileMD is available at github.com/finnlueth/ProtProfileMD. The associated training data and model weights are available at huggingface.co/datasets/finnlueth/ProtProfileMD and huggingface.co/finnlueth/ProtProfileMD.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.