Back

Bioinformatic process for the identification and characterization of bacterial repeat-in-toxin adhesins

Hansen, T.; Graham, L. A.; Soares, B. P.; Lee, D.; Gagnon, J. R.; Dykstra-MacPherson, T.; Guo, S.; Davies, P. L.

2025-09-30 biochemistry
10.1101/2025.09.30.679566 bioRxiv
Show abstract

Gram-negative bacteria attach to host surfaces using ligand-binding domains (LBDs) at the distal tips of fibrillar RTX adhesins. Blocking the initial binding interaction(s) can potentially prevent colonization and subsequent biofilm formation and infection. To this end, adhesins must be identified, and it is also essential to determine the dominant adhesin type for those species that have more than one. RTX adhesins are frequently the largest proteins within each species (ranging from 1500 to 15,000 aa) and are often misannotated as incomplete/pseudogene products because their highly repetitive nature confounds genome assemblies from short-read technologies. Our bioinformatic process collates predicted proteins from long-read assemblies, which are then clustered based on the similarity of their C-terminal regions where the LBDs are typically located. RTX adhesins are identified by their length and domain structure and are modelled using AlphaFold3. An exhaustive search of multiple strains from seven species revealed a total of 35 different RTX adhesins that map to 16 different loci, with differing arrangements of LBDs that include putative carbohydrate-binding modules and von Willebrand Factor A-like domains. Notably, similar adhesins are sometimes found in multiple species, either by descent or through DNA uptake, and three species have an RTX adhesin of uncertain function because it lacks an obvious LBD. ImportanceMany bacteria initiate infection by reaching out with large, complex proteins called adhesins to attach themselves to a host cell. The DNA sequences of adhesins are difficult to read, due to their length and repetitiveness. This study leverages recent technological advances like "long-read sequencing" and structural modelling to identify and characterize the adhesins in seven species of harmful bacteria: Acinetobacter baumannii, Aeromonas hydrophila, Aeromonas salmonicida, Bordetella parapertussis, Legionella pneumophila, Vibrio parahaemolyticus, and Vibrio vulnificus. We identified thirty-five unique versions of adhesins and demonstrated their mix-and-match architecture. This research provides a foundation for strategies to block bacteria from binding surfaces, offering a vital alternative treatment as antibiotic resistance continues to rise.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.