ProDive reveals pervasive cross-family protein fragment reuse
Chen, X.; Tian, P.
Show abstract
Cross-family reuse of short protein fragments has been a long-standing mystery whose resolution first demands an algorithm for their systematic detection. Here we introduce ProDive, a closed-form symmetric KL divergence between profile HMMs that enables GPU-accelerated, fragment-level screening across all 25,545 Pfam families. ProDive identifies [~]318,000 cross-family fragment correspondences involving compact cores of 8-13 residues with RMSD values far below random background. Their organisation into diverse graph communities and four-fold enrichment in de novo designed proteins point away from family-specific functions and toward a general biophysical property. Their helix dominance and moderate solvent exposure suggest a role in folding initiation--a link corroborated by overlap with experimentally measured{phi} -values and by a monotonic density gradient across disordered regions. Together, these observations converge on a single explanation: cross-family fragment reuse likely reflects shared requirements for early structure formation during folding, the one biophysical constraint common to all proteins.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep-Learning Structure Elucidation from Single-Mutant Deep Mutational Scanning 95%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 95%
- Improved protein structure refinement guided by deep learning based accuracy estimation 95%
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- The amino acid sequence determines protein abundance through its conformational stability and reduced synthesis cost. 95%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.