SLAE: Strictly Local All-atom Environment for Protein Representation
Chen, Y.; Lu, T.; Zhao, C.; Wayment-Steele, H. K.; Huang, P.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWBuilding physically grounded protein representations is central to computational biology, yet most existing approaches rely on sequence-pretrained language models or backbone-only graphs that overlook side-chain geometry and chemical detail. We present SLAE, a unified all-atom framework for learning protein representations from each residues local atomic neighborhood using only atom types and interatomic geometries. To encourage expressive feature extraction, we introduce a novel multi-task autoencoder objective that combines coordinate reconstruction, sequence recovery, and energy regression. SLAE reconstructs allatom structures with high fidelity from latent residue environments and achieves state-of-the-art performance across diverse downstream tasks via transfer learning. SLAEs latent space is chemically informative and environmentally sensitive, enabling quantitative assessment of structural qualities and smooth interpolation between conformations at all-atom resolution.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProtMamba: a homology-aware but alignment-free protein state space model 97%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 96%
- Cross-Modality and Self-Supervised Protein Embedding for Compound-Protein Affinity and Contact Prediction 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- Cracking the black box of deep sequence-based protein-protein interaction prediction 94%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.