Back

Functional diversity across families of bacterial metalloregulators: what can we learn about specificity from sequence similarity?

Rondon, J. J.; Antelo, G. T.; Capdevila, D. A.

2026-05-05 bioinformatics
10.64898/2026.05.01.721428 bioRxiv
Show abstract

The exponential growth of sequence databases, driven by large-scale genome sequencing, has created a major challenge for the functional annotation of proteins, particularly within highly divergent families such as bacterial metal-responsive transcription factors. In these systems, low sequence identity places many proteins within the so-called "twilight zone" of the proteome, where homology inference and functional assignment become unreliable. Here, we integrate sequence similarity networks (SSNs) with structural approaches to explore the functional diversity of twelve metalloregulatory families. Using SSNs, we partition each family into putative isofunctional clusters and map available experimental annotations onto network topology, revealing substantial heterogeneity in both sequence diversity and functional characterization across families. While some families, such as ArsR and CsoR, display relatively well-defined functional landscapes, others, including LysR, TetR, and GntR, remain largely unexplored despite their large sequence space. Structural comparisons further show that, despite extensive sequence divergence, conserved architectural features underpin DNA recognition and regulatory mechanisms across families. Focusing on the MerR and Fur families, we identify conserved residues associated with inducer binding and DNA recognition, and uncover distinct functional subgroups, including metal-specific sensors and regulators with alternative signaling mechanisms. Finally, we demonstrate that cluster-derived HMM profiles enable sensitive detection of candidate regulators in non-model genomes, revealing lineage-specific expansions and diversification of metal-sensing repertoires. Together, our results provide a framework for mapping functional diversity in highly divergent protein families and highlight the potential of combining SSNs and structural information to guide the discovery of novel transcriptional sensors.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.