Back

Enhanced Protein-Protein Interaction Discovery via AlphaFold-Multimer

Kim, A.-R.; Hu, Y.; Comjean, A.; Rodiger, J.; Mohr, S. E.; Perrimon, N.

2024-02-21 bioinformatics
10.1101/2024.02.19.580970 bioRxiv
Show abstract

Accurately mapping protein-protein interactions (PPIs) is critical for elucidating cellular functions and has significant implications for health and disease. Conventional experimental approaches, while foundational, often fall short in capturing direct, dynamic interactions, especially those with transient or small interfaces. Our study leverages AlphaFold-Multimer (AFM) to re-evaluate high-confidence PPI datasets from Drosophila and human. Our analysis uncovers a significant limitation of the AFM-derived interface pTM (ipTM) metric, which, while reflective of structural integrity, can miss physiologically relevant interactions at small interfaces or within flexible regions. To bridge this gap, we introduce the Local Interaction Score (LIS), derived from AFMs Predicted Aligned Error (PAE), focusing on areas with low PAE values, indicative of the high confidence in interaction predictions. The LIS method demonstrates enhanced sensitivity in detecting PPIs, particularly among those that involve flexible and small interfaces. By applying LIS to large-scale Drosophila datasets, we enhance the detection of direct interactions. Moreover, we present FlyPredictome, an online platform that integrates our AFM-based predictions with additional information such as gene expression correlations and subcellular localization predictions. This study not only improves upon AFMs utility in PPI prediction but also highlights the potential of computational methods to complement and enhance experimental approaches in the identification of PPI networks.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.