Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction
Xiao, W.; Zhang, Y.; Gong, C.; Zhang, H.; Ma, W.; Liu, Z.; Chen, X.; Guan, J.; Wang, L.
Show abstract
We introduce Protenix-v1 (PX-v1), the first fully open-source structure prediction model to attain superior performance to AlphaFold3 while strictly adhering to the same training data cutoff, model size, and inference budget. Beyond standard evaluations, we highlight the effectiveness of inference-time scaling behavior of Protenix-v1, demonstrating that increasing the sampling budget yields consistent improvements in prediction quality--a behavior previously observed in AlphaFold3 and largely absent from prior open-source models. In addition to improved accuracy, Protenix-v1 incorporates key capabilities including protein template integration and RNA MSA support. Furthermore, to better support real-world applications such as drug discovery, we additionally release Protenix-v1-20250630, a variant trained on a larger dataset (cutoff: June 30, 2025), delivering further improved prediction accuracy. Finally, we identify limitations in existing benchmarking practices and provide updated evaluation tools and year-stratified benchmarks to support more reliable and transparent assessment. Collectively, these contributions provide a robust foundation for the Protenix series and the broader field.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mapping the space of protein binding sites with sequence-based protein language models 96%
- FlowPacker: Protein side-chain packing with torsional flow matching 96%
- QDeep: distance-based protein model quality estimation by residue-level ensemble error classifications using stacked deep residual neural networks 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.