A Structural Antibody Benchmark of AlphaFold3 reveals Hallucinated Epitopes and a Bias for Orderness
Solanki, A.; Maurya, N. S.; Ramlakhan, M.; Li, R.; Chen, W.; Wu, Z.; Zheng, W. J.
Show abstract
AlphaFold3 has shown promise as a tool for predicting antibody-antigen binding, yet its performance across large datasets has not been fully characterized. In this study, 3401 experimentally validated antibody-antigen complexes were sourced from the Structural Antibody Database and screened alongside 23798 negative controls to benchmark AlphaFold3s binding prediction capabilities. Confidence metrics including Predicted Aligned Error and Interface Predicted Template Modeling score were used to achieving a maximum recall of 53% at 100 inference seeds. Several factors were found to influence prediction accuracy: a notable bias was observed toward antibodies derived from X-ray crystallography structures versus those from electron microscopy, and positive prediction rates were found to decrease with increasing target protein size and surface area. In contrast, neither the amino acid composition or lengths of the complementarity determining regions, nor training data leakage were found to introduce significant bias. An innate false positive rate of approximately 3% was identified, with AF3 shown to hallucinate plausible binding interfaces across the surface of decoy targets while avoiding disordered regions. Epitope mapping using DockQ, epitope shift, and antibody displacement revealed that approximately 34% of false negatives retained the correct epitope location despite poor structural alignment, suggesting that conformation refinement tools could recover additional true binding predictions. These findings provide a comprehensive characterization of AlphaFold3s strengths and limitations for antibody screening in computational drug discovery. Key MessagesO_LIAlphaFold3 has a recall of 50% and an innate false positive prediction rate of 3%. C_LIO_LIFalse negative predictions can still feature the correct epitope despite poor RMSD. C_LIO_LIFactors such as disorder and target size impact accuracy. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 94%
- CompScore: boosting structure-based virtual screening performance by incorporating docking scoring functions components into consensus scoring 94%
- CLIMBS: assessing Carbohydrate-Protein interactions through a graph neural network classifier using synthetic negative data 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.