Back

UdonPred: Untangling Protein Intrinsic Disorder Prediction

Schlensok, J.; Wagemann, D.; Senoner, T.; Haak, M.; Rost, B.

2026-01-27 bioinformatics
10.64898/2026.01.26.701679 bioRxiv
Show abstract

MotivationRegions in intrinsic disordered proteins (IDPs) constitute important continuous aspects of protein function. While their existence on a structural continuum is widely accepted, most computational predictions have, nevertheless, focused on binary classifications. Existing datasets are severely limited in size and experimental evidence for continuous disorder. ResultsBuilding on recently released datasets of continuous protein disorder and flexibility, we introduce UdonPred, a lightweight neural network exclusively inputting embeddings from the protein Language Model (pLM) ProstT5 to predict per-residue protein disorder from sequence alone. Training and evaluating UdonPred on seven datasets with divergent definitions of disorder and flexibility suggests that not model capacity, but agreement and nuance of disorder annotations, remains the main driver of performance. Binary disorder annotations can be reliably predicted from a multitude of different disorder and flexibility datasets, but there is still room for improvement in predicting continuous disorder. AvailabilityAll code and data used for training and evaluation is available under an open-source license at https://github.com/davidwagemann/udonpred. Contactassistant@rostlab.org Supplementary informationSupplementary data are available at Journal Name online.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.