Pandoomain, a scalable pipeline for genomic and protein domain context analysis, reveals widespread PT-TG domain architectural diversity and novel polymorphic toxins
Soto, E. B.; Oliver, A. J.; de Moraes, M. H.
Show abstract
The rapid expansion of bacterial genome databases presents significant opportunities for functional discovery, yet a large fraction of genes and protein domains remain uncharacterized. Analyzing genomic context and domain architecture are powerful approaches for functional inference, but existing tools often lack the scalability and integrated workflow required for high-throughput analysis. To address this, we developed Pandoomain, a Snakemake pipeline that automates the acquisition of genomes from NCBI, identifies proteins of interest using Hidden Markov Models (HMMs), and performs systematic domain annotation and gene neighborhood analysis. We demonstrate the utility of Pandoomain through a comprehensive analysis of the poorly characterized pre-toxin TG (PT-TG) domain across 347,289 bacterial genomes. Our analysis revealed 10,226 PT-TG-containing proteins organized into 312 unique domain architectures, highlighting their association with diverse interbacterial antagonistic systems, including the Type VI, Type VII, and CDI systems. By leveraging genomic context, we identified a novel variant of the WXG trafficking domain, termed W10XG, and subsequently discovered 24 new families of associated toxin domains. We experimentally validated six of these toxins, confirming that five are neutralized by their cognate immunity proteins. Furthermore, our analysis revealed a significant enrichment of mobile genetic elements near W10XG and WXG domains compared to other trafficking domains, suggesting these loci are hotspots for genomic diversification. Pandoomain is an accessible tool that enables systematic, large-scale exploration of protein domains, and our analysis of the PT-TG domain provides a rich resource for future investigations into the mechanisms and evolution of bacterial antagonism.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Klebsiella pneumoniae type VI secretion system-mediated microbial competition is PhoPQ controlled and reactive oxygen species dependent 95%
- The fitness landscape of the African Salmonella Typhimurium ST313 strain D23580 reveals unique properties of the pBT1 plasmid 94%
- Discovery and characterization of a Gram-positive Pel polysaccharide biosynthetic gene cluster 94%
Similar papers in this journal
Similar papers in this journal
- Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp. 94%
- cazy_webscraper: local compilation and interrogation of comprehensive CAZyme datasets 94%
- Taxonomic distribution of SbmA/BacA and BacA-like antimicrobial peptide transporters suggests independent recruitment and convergent evolution in host-microbe interactions 94%
Similar papers in this journal
- Bacterial hemophilin homologs and their specific type eleven secretor proteins have conserved roles in heme capture and are diversifying as a family 95%
- Large Scale Discovery of Microbial Fibrillar Adhesins and Identification of Novel Members of Adhesive Domain Families 94%
- Eight Unexpected Selenoprotein Families in ABC transport, in Organometallic Biochemistry in Clostridium difficile and other anaerobes, and in Methylmercury Biosynthesis. 94%
Similar papers in this journal
- Prediction of Burkholderia pseudomallei DsbA substrates identifies potential virulence factors and vaccine targets 94%
- Identification of the Clostridial cellulose synthase and characterization of the cognate glycosyl hydrolase, CcsZ 94%
- Gene co-expression network analysis of the human gut commensal bacterium Faecalibacterium prausnitzii based on WGCNA in R-Shiny 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.