InteracTor: A new integrative feature extraction toolkit for improved characterization of protein structural properties
Silva, J. C. F.; Schuster, L.; Sexson, N.; Kirst, M.; Resende, M. F. R.; Dias, R.
10.1101/2024.10.07.616705 bioRxivShow abstract
Understanding the structural and functional diversity of protein families is crucial for elucidating their biological roles. Traditional analyses often focus on primary and secondary structures, which include amino acid sequences and local folding patterns like alpha helices and beta sheets. However, primary and secondary structures alone may not fully represent the complex interactions within proteins. To address this limitation, we developed a new algorithm (InteracTor) to analyze proteins by extracting features from their three-dimensional (3D) structures. The toolkit extracts interatomic interaction features such as hydrogen bonds, van der Waals interactions, and hydrophobic contacts, which are crucial for understanding protein dynamics, structure, and function. Incorporating 3D structural data and interatomic interaction features provides a more comprehensive understanding of protein structure and function, potentially enhancing downstream predictive modeling capabilities. By using the extracted features in Mutual Information scoring (MI), Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), and hierarchical clustering analysis as use cases, we identified clear separations among protein structural families, highlighting distinct functional aspects. Our analysis revealed that interatomic interaction features were more informative than protein secondary structure features, providing insights into potential structural and functional properties. These findings underscore the significance of considering tertiary structure in protein analysis, offering a robust framework for future studies aiming at enhancing the capabilities of models for protein function prediction and drug discovery.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 96%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
Similar papers in this journal
Similar papers in this journal
- The blobulator: a toolkit for identification and visual exploration of hydrophobic modularity in protein sequences 96%
- BINANA 2.0: Characterizing Protein/Ligand Interactions in Python and JavaScript 95%
- ANABAG: Annotated Antibody Antigen dataset with unique features for Antibody Engineering Applications 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.