The GenPPI tool enhanced Protein Interaction Network Generation with Machine Learning-Based Protein Similarity Inference
William, A.; Godoy, I.; Marquez, C.; Silva, L.; Prado, M.; Avila, N.; Santos, A. R. d.
Show abstract
AbstractO_ST_ABSBackgroundC_ST_ABSComputational prediction of protein-protein interactions (PPIs) is crucial for understanding cell biology and drug development, offering an alternative to costly experimental methods. The original GenPPi software advanced ab initio PPI network prediction from bacterial genomes but was limited by its reliance on high sequence similarity. This work introduces GenPPi 1.5 to enhance these predictive capabilities. ResultsGenPPi 1.5 incorporates a Random Forest (RF) algorithm, trained on 60 biophysical features from amino acid propensity indices, to classify protein similarity even in low sequence identity scenarios (targeting >65% identity). To manage computational complexity from the increased interactions generated by the RF model, especially in extensive conserved phylogenetic profiles, we developed and integrated the Reduced Interaction Sampling (RIS) algorithm. RIS stochastically samples interactions within these profiles, optimizing performance for complete genome analysis. Extensive simulations across various configurations validated the methodology. RF integration significantly broadened GenPPis predictive power; application to Buchnera aphidicola showed up to 62% overlap with STRING database interactions. Analysis of RIS demonstrated that while introducing some randomness, critical node identification remains robust, particularly for Top N values[≥] 100, indicating minimal compromise to network integrity. ConclusionThe combination of Machine Learning (RF) and the RIS algorithm in GenPPi 1.5 represents a significant advancement. It overcomes the highsimilarity dependency of the previous version while efficiently handling complex genomes. GenPPi 1.5 provides a robust and scalable alignment-free PPI prediction solution, enabling users to train custom models tailored to specific genomic contexts. GenPPi is freely available on our website https://genppi.facom.ufu.br/, its source code is hosted on GitHub https://github.com/santosardr/genppi, and it can be easily installed via the Python Package Index using the command pip install genppipy.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 95%
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 95%
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 95%
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 96%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 94%
- MLDSP-GUI: An alignment-free standalone tool with an interactive graphical user interface for DNA sequence comparison and analysis 94%
Similar papers in this journal
Similar papers in this journal
- Structome-TM: Complementing dataset assembly for structural phylogenetics by addressing size-based biases 94%
- Genomic style: yet another deep-learning approach to characterize bacterial genome sequences 94%
- Accelerating Protein-Protein Interaction screens with reduced AlphaFold-Multimer sampling 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.