Unbiased learning of protein conformational representation via unsupervised random forest
Sahil, M.; Ahalawat, N.; Mondal, J.
Show abstract
Accurate data representation is paramount in biophysics to capture the functionally relevant motions of biomolecules. Traditional feature selection methods, while effective, often rely on labeled data based on prior knowledge and user-supervision, limiting their applicability to novel systems. Here, we present unsupervised random forest (URF), a self-supervised adaptation of traditional random forests that identifies functionally critical features of biomolecules without requiring prior labels. By devising a memory-efficient implementation, we first demonstrate URFs capability to learn important sets of inter-residue features of a protein and subsequently to resolve its complex conformational landscape, performing at par or surpassing its traditional supervised counterpart and 15 other leading baseline methods. Crucially, URF is supplemented by an internal metric, the learning coefficient, which automates the process of hyper-parameter optimization, making the method robust and user-friendly. URFs remarkable ability to learn important protein features in an unbiased fashion was validated against 10 independent protein systems including both both folded and intrinsically disordered states. In particular, benchmarking investigations showed that the protein representations identified by URF are functionally meaningful in comparison to current state-of-the-art deep learning methods. As an application, we show that URF can be seamlessly integrated with downstream analyses pipeline such as Markov state models to attain better resolved outputs. The investigation presented here establishes URF as a leading tool for unsupervised representation learning in protein biophysics.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Decoding protein-membrane binding interfaces from surface-fingerprint-based geometric deep learning and molecular dynamics simulations 97%
- Covalent adducts formed by the androgen receptor transactivation domain and small molecule drugs remain disordered 97%
- Learning Binding Affinities via Fine-tuning of Protein and Ligand Language Models 96%
Similar papers in this journal
- Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction 96%
- DeepRank: A deep learning framework for data mining 3D protein-protein interfaces 96%
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 96%
Similar papers in this journal
- Folding-upon-binding pathways of an intrinsically disordered protein from a deep Markov state model 97%
- Fast calculation of small-angle scattering profiles of dense protein solutions modeled at the all-atom level 96%
- Prediction of phase separation propensities of disordered proteins from sequence 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.