Reliable Identification of Homodimers Using AlphaFold
Danielsson, S. N.; Elofsson, A.
Show abstract
MotivationProtein-protein interactions are central for understanding biological processes. The ability to predict interaction partners is extremely valuable for avoiding costly, time-consuming experiments. It has been shown that AlphaFold has an unsurpassed ability to accurately evaluate interacting protein pairs. However, a protein can also form homomeric interactions, i.e. interact with itself. ResultsWe found that AlphaFold yielded a significantly higher false-positive rate for identifying homodimers than for heterodimers. True Positive Rate (TPR) at 1% False Positive Rate (FPR) drops from 63% for heterodimers to 18% for homodimers. When we investigated the high-scoring false positives, i.e., non-homodimers with high AlphaFold scores when predicted as such, we found that their homologs were enriched for homomultimeric proteins. Using a simple logistic regression model that combines AlphaFold scores with structural and homology information, we increased the TPR (at 1% FPR) to 42{+/-}8% (5-fold cross-validation) from 19%. If we excluded the homology information, we achieved a TPR of 28{+/-}7%, which is still better than using AlphaFold metrics. Availability and implementationAll data are available from Zenodo DOI:10.5281/zenodo.17738668 and all code from https://github.com/SarahND97/alphafold-homodimers. Supplementary informationSupplementary information is available online.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 95%
- Improved protein complex prediction with AlphaFold-multimer by denoising the MSA profile 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.