P(all-atom) Is Unlocking New Path For Protein Design
Qu, W.; Guan, J.; Ma, R.; Zhai, K.; Wu, W.; Wang, H.
Show abstract
We introduce Pallatom, an innovative protein generation model capable of producing protein structures with all-atom coordinates. Pallatom directly learns and models the joint distribution P(structure, seq) by focusing on P(all-atom), effectively addressing the interdependence between sequence and structure in protein generation. To achieve this, we propose a novel network architecture specifically designed for all-atom protein generation. Our model employs a dual-track framework that tokenizes proteins into residue-level and atomic-level representations, integrating them through a multi-layer decoding process with "traversing" representations and recycling mechanism. We also introduce the atom14 representation method, which unifies the description of unknown side-chain coordinates, ensuring high fidelity between the generated all-atom conformation and its physical structure. Experimental results demonstrate that Pallatom excels in key metrics of protein design, including designability, diversity, and novelty, showing significant improvements across the board. Our model not only enhances the accuracy of protein generation but also exhibits excellent sampling efficiency, paving the way for future applications in larger and more complex systems.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FlowPacker: Protein side-chain packing with torsional flow matching 97%
- From Atoms to Fragments: A Coarse Representation for Functional and Efficient Protein Design 97%
- QDeep: distance-based protein model quality estimation by residue-level ensemble error classifications using stacked deep residual neural networks 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.