Back

Efficient Enumeration and Visualization of Helix-coil Ensembles

Schmidler, S. C.; Hughes, R. G.; Oas, T. G.; Zhao, S.

2023-09-17 biophysics
10.1101/2023.09.16.558052 bioRxiv
Show abstract

Helix-coil models are routinely used to interpret CD data of helical peptides or predict the helicity of naturally-occurring and designed polypeptides. However, a helix-coil model contains significantly more information than mean helicity alone, as it defines the entire ensemble - the equilibrium population of every possible helix-coil configuration - for a given sequence. Many desirable quantities of this ensemble are either not obtained as ensemble averages, or are not available using standard helicity-averaging calculations. Enumeration of the entire ensemble can allow calculation of a wider set of ensemble properties, but the exponential size of the configuration space typically renders this intractable. We present an algorithm that efficiently approximates the helix-coil ensemble to arbitrary accuracy, by sequentially generating a list of the M highest populated configurations in descending order of population. Truncating this list of (configuration, population) pairs at a desired accuracy provides an approximating sub-ensemble. We demonstrate several uses of this approach for providing insight into helix-coil ensembles and folding mechanisms, including landscape visualization. O_TEXTBOXSIGNIFICANCE Helix-coil models define the probability distribution of helix-coil configurations for a polypeptide (a helix-coil ensemble). Each configuration specifies which residues are -helical and which are not. We used an accurate helix-coil model, paired with concepts from the field of probabilistic graphical modeling, to devise an algorithm capable of enumerating helix-coil configurations in order of decreasing probability. By enumerating ensembles for a representative set of peptides we find that helix-coil ensembles tend to be highly concentrated, with the vast majority of probability mass assigned to a relatively small set of configurations from a configuration space that is often astronomical in size. This result facilitates the development of new and accurate methods for analyzing and predicting the behavior of helical polypeptides. C_TEXTBOX

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.