Fold first, ask later: structure-informed function annotation of Pseudomonas phage proteins
Longin, H.; Bouras, G.; Grigson, S. R.; Edwards, R. A.; Hendrix, H.; Lavigne, R.; van Noort, V.
Show abstract
Phages, the viruses of bacteria, harbor an incredibly diverse repertoire of proteins capable of manipulating their bacterial hosts, inspiring many medical and biotechnological applications. However, to date, only a limited subset of that repertoire can be exploited, due to the difficulties in functionally elucidating these proteins. In this study, we investigated several structure-informed approaches to annotate hypothetical proteins from Pseudomonas infecting phages. We curated a representative dataset of over 10,000 proteins derived from NCBI, for which we predicted protein structures with ColabFold and assessed structural similarity via FoldSeek against the PDB, AlphaFold, and Phold databases. We evaluated multiple annotation strategies, including sequence-based (Pharokka), and structure-based (FoldSeek, Phold) methods. Our results show that up to 43 % of truly unannotated proteins can be functionally annotated when combining structure-informed approaches with UniProt-derived annotations. We highlight the complementarity of different databases and the importance of annotation quality filtering. This work provides a valuable resource of predicted structures and annotations, and offers insights into optimizing structure-based annotation pipelines for viral proteins, paving the way for deeper exploration of phage biology and its applications.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Do Newly Born Orphan Proteins Resemble Never Born Proteins? A Study Using Three Deep Learning Algorithms 95%
- Generalizable strategy to analyze domains in the context of parent protein architecture: A CheW case study 94%
- Prediction of protein assemblies by structure sampling followed by interface-focused scoring 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.