Back

ppigFinder: an integrated desktop application for bacterial genome annotation and AlphaFold 3 based protein protein interaction screening

Oka, G. U.; Adan, W. C.; Calomeno, C. Q.; de Souza, R. F.

2026-08-26 bioinformatics
10.64898/2026.08.23.746524 bioRxiv
Show abstract

Motivation. AlphaFold-based structure prediction has transformed structural biology by enabling accurate protein modelling and providing a powerful framework for inferring protein-protein interactions (PPIs). However, discovering candidate PPIs directly from genome sequences remains a fragmented and largely trial-and-error process, typically requiring separate tools for open reading frame (ORF) prediction, functional annotation, candidate selection, iterative testing of potential partners, manual preparation of individual structural-prediction jobs, and downstream interpretation of confidence metrics. Results. We present Protein-Protein Interaction Genomic Finder (ppigFinder), a standalone, cross-platform desktop application that integrates these steps into a project-oriented graphical workflow for genome-based PPI discovery from nucleotide sequence data. ppigFinder combines ORF prediction, functional annotation, genomic-neighbourhood inspection, AlphaFold 3 job generation, remote job submission, and structural-confidence analysis within a single environment. As a proof of concept, we performed a VirD4-centered AlphaFold 3 interactome screen in Xanthomonas citri pv. citri strain 306, modelling VirD4 (ORF2601) against all 4,303 predicted chromosomal ORFs. Ranking by the minimum interchain predicted aligned error (PAE_min) placed all 14 XVIPCD-containing effector candidates within the top 1% of predictions, with the six top-ranked models corresponding to XVIP candidates. The screen also recovered an XVIPCD-containing protein absent from the reference genome annotation and identified high-confidence candidates predicted to bind VirD4 at a surface opposite to the XVIPCD-binding site.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.