Back

Improving Inverse Folding models at Protein Stability Prediction without additional Training or Data

Dutton, O.; Bottaro, S.; Invernizzi, M.; Redl, I.; Chung, A.; Hoffmann, F.; Henderson, L.; Ruschetta, S.; Airoldi, F.; Owens, B. M.; Foerch, P.; Fisicaro, C.; Tamiola, K.

2024-09-09 biophysics
10.1101/2024.06.15.599145 bioRxiv
Show abstract

Deep learning protein sequence models have shown outstanding performance at de novo protein design and variant effect prediction. We substantially improve performance without further training or use of additional experimental data by introducing a second term derived from the models themselves which align outputs for the task of stability prediction. On a task to predict variants which increase protein stability the absolute success probabilities of PO_SCPLOWROTEINC_SCPLOWMPNN and ESMO_SCPLOWIFC_SCPLOW are improved by 11% and 5% respectively. We term these models PO_SCPLOWROTEINC_SCPLOWMPNN-O_SCPLOWDDC_SCPLOWG and ESMO_SCPLOWIFC_SCPLOW-O_SCPLOWDDC_SCPLOWG.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.