Back

AF-CALVADOS: AlphaFold-guided simulations of multi-domain proteins at the proteome level

von Bülow, S.; Johansson, K. E.; Lindorff-Larsen, K.

2025-12-22 biophysics
10.1101/2025.10.19.683306 bioRxiv
Show abstract

Deep-learning methods have transformed our ability to predict the three-dimensional structures of folded proteins from sequence, and coarse-grained simulations have made it possible to study intrinsically disordered proteins at the proteome scale. More than half of human proteins, however, contain mixtures of disordered regions and one or more folded domains, and the biological function of these multi-domain proteins depends on the interplay between the folded and disordered regions. Here, we developed AF-CALVADOS, a coarse-grained simulation model that is informed by AlphaFold to model the dynamics of intrinsically disordered proteins and multi-domain proteins containing mixtures of folded and disordered regions. AF-CALVADOS leverages information from AlphaFold 2 to model folded regions that we then integrate with the coarse-grained CALVADOS model. Our automated framework makes it possible to perform simulations of any soluble folded or disordered protein without manually defining the folded regions, enabling scaling to the proteome level. We validate AF-CALVADOS using experimental SAXS data for more than 400 proteins and find that it performs well across proteins with varying amounts of ordered and disordered regions. We demonstrate the scalability of AF-CALVADOS by performing simulations of 12,483 intracellular human proteins and make the data freely available; we envisage that large-scale simulation data generated by AF-CALVADOS can be used to benchmark or train machine learning models for flexible, multi-domain proteins. The conformational ensembles can also be used to study sequence-dynamics-function relationships at scale, and can shed light on the interplay between folded and disordered regions. We exemplify this by analysing the disordered regions in 1,487 human transcription factors.

Published in Protein Science (predicted rank #4) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.