Back

From Dataset Curation to Unified Evaluation: Revisiting Structure Prediction Benchmarks with PXMeter

Ma, W.; Liu, Z.; Yang, J.; Lu, C.; Zhang, H.; Xiao, W.

2025-07-22 bioinformatics
10.1101/2025.07.17.664878 bioRxiv
Show abstract

Recent advances in deep learning have significantly improved the accuracy of structure prediction for biomolecular complexes; however, robust evaluation of these models remains a major challenge. We introduce PXMeter, an open-source toolkit that support consistent and reproducible evaluation of diverse predictive models across a broad spectrum of biological complex structures. PXMeter provides a unified and reproducible benchmarking framework, offering valuable insights to support the ongoing improvement of structure prediction methods. We also present a high-quality benchmark dataset curated from recently deposited structures in the Protein Data Bank (PDB). These entries are manually reviewed to exclude non-biological interactions, ensuring reliable evaluation. Using these resources, we conducted a comprehensive benchmark of several structure prediction models, namely Chai-1, Boltz-1, and Protenix. Our benchmarking results demonstrate the advancements achieved by deep learning models, while also identifying ongoing challenges--especially in modeling protein-protein and protein-RNA interactions. Project Pagehttps://github.com/bytedance/PXMeter

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.