Back

Defining the ESKAPE pathogen prophage repertoire with PHORAGER

Dyball, X.; Ponsero, A. J.; Docherty, J. A. D.; Telatin, A.; Crost, E. H.; Juge, N.; Cook, R.; Adriaenssens, E. M.

2026-08-05 bioinformatics
10.64898/2026.08.05.742953 bioRxiv
Show abstract

Prophages are major drivers of bacterial evolution, mediating horizontal gene transfer and lysogenic conversion to alter host phenotypes. Nevertheless, identifying prophages within bacterial genomes remains challenging due to their heterogeneity and similarity to other mobile genetic elements. Here we present PHORAGER (Prophage Hunting, vOtu Retrieval, Annotation and Genomic ExploRation), a scalable Nextflow pipeline for the standardised identification and quality assessment of prophages from bacterial genomes. PHORAGER incorporates bacterial genome pre-processing, consolidation of predictions from multiple mining tools, annotation-based filtering to reduce false positives, and generation of ready-to-analyse summary tables. We validated PHORAGER using 30,824 publicly available ESKAPE pathogen genomes. PHORAGER recovered more high-quality prophages than individual mining tools alone, and through extensive quality assessments removed a substantial number of false-positive predictions. In total 23,132 putative prophages were identified, the majority belonging to the class Caudoviricetes, and exhibiting a high degree of host-specificity. Putative antimicrobial resistance genes were detected in 0.48% of prophages, whereas virulence factors were most abundant in S. aureus prophages. ESKAPE prophages also frequently encoded anti-phage defence systems. PHORAGER is freely available as open-source software and the ESKAPE prophage collection generated in this study provides a reusable resource for further investigations. GRAPHICAL ABTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/742953v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@1ce582org.highwire.dtl.DTLVardef@11fc004org.highwire.dtl.DTLVardef@17765f5org.highwire.dtl.DTLVardef@1c6f0ff_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.